ContentsThe library

Judging a System, Not a Model

Three Points on Four Hundred Cases Is Nothing

Last timeLetting a Model Mark It

A score without an error bar cannot answer the only question anybody asks of it. Here is the arithmetic, and the smallest difference your set can honestly detect.

The error on one score

Your evaluation produces 74 per cent. Last week it produced 71. Did anything

improve?

The question cannot be answered from those two numbers, and the reason is that

a score is a proportion measured on a sample. Run the same unchanged system

against a different sample of the same size and you get a different number.

That spread is not a defect in the evaluation, it is what measuring a sample

means, and the only mistake available is to ignore it.

FIG 1The uncertainty on a score
the standard error: roughly how far a repeat measurement would typically land from this one
the score, as a fraction between 0 and 1
the number of cases in the set
Two things to notice. The top is largest at a score of one half and falls toward the extremes, so a system scoring 95 per cent is measured more precisely than one scoring 50. And the bottom is a square root, which is why making an evaluation twice as precise costs four times as many cases.

At 400 cases and a score near 70 per cent, the error is about 2.3 points. So a

single score of 74 is really somewhere around 72 to 76, and a move from 71 to

74 is well inside the range you would see from a system that did not change at

all.

FIG 2The same unchanged system, four times
stepruncasesscorewhat changed in the systemwhat happened
114000.71nothingThe first measurement becomes the baseline simply because it happened first.
224000.74nothingThree points up. If this had followed a prompt edit, the edit would now be considered an improvement and kept.
334000.69nothingAnd back down, five points below the previous run. Still nothing changed.
444000.73nothingThe spread across four runs is four points, all of it noise. Any explanation offered for any of these moves would be a story about nothing.
4 steps
Resampling the set from the same traffic each time, with the system frozen. This is the single most useful experiment to run before trusting an evaluation, because it shows the team what a change of nothing looks like.

The error on a difference

Now the harder part, which is where most reporting goes wrong. You do not

usually care about one score; you care about the difference between two.

Two independently measured scores each carry their own error, and when you

subtract them the errors do not cancel. They combine: the error on the

difference is the square root of the sum of the squares of the two errors. For

two equal-sized sets that is about 1.4 times the error on either one.

So at 400 cases, with an error of 2.3 points on each score, the error on the

difference is about 3.3 points. A three-point difference is therefore

indistinguishable from zero, and to call it a result you want the difference to

be roughly twice its error.

The lesson stops here

3 more paragraphs to go

You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.

The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.

See the planThe contents

This is the reading half

Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.

The contents

The rest of this course

  1. 01The Boundary Decides What the Numbers Mean
  2. 02The Set Is a Sample, and You Chose the Samplingopening only
  3. 03If Two People Score It Differently, You Have No Measurementopening only
  4. 04A Marker That Never Gets Tired and Has Four Known Habitsopening only
  5. 05Three Points on Four Hundred Cases Is Nothingyou are here
  6. 06Run Both on the Same Cases and Most of the Noise Disappearsopening only
  7. 07A Small Fast Set That Blocks, and Two Slower Ones That Do Notopening only
  8. 08Four Reasons Your Number Disagrees With Your Usersopening only

Read alongside