Three Points on Four Hundred Cases Is Nothing
Last timeLetting a Model Mark It
A score without an error bar cannot answer the only question anybody asks of it. Here is the arithmetic, and the smallest difference your set can honestly detect.
The error on one score
Your evaluation produces 74 per cent. Last week it produced 71. Did anything
improve?
The question cannot be answered from those two numbers, and the reason is that
a score is a proportion measured on a sample. Run the same unchanged system
against a different sample of the same size and you get a different number.
That spread is not a defect in the evaluation, it is what measuring a sample
means, and the only mistake available is to ignore it.
- the standard error: roughly how far a repeat measurement would typically land from this one
- the score, as a fraction between 0 and 1
- the number of cases in the set
At 400 cases and a score near 70 per cent, the error is about 2.3 points. So a
single score of 74 is really somewhere around 72 to 76, and a move from 71 to
74 is well inside the range you would see from a system that did not change at
all.
| step | run | cases | score | what changed in the system | what happened |
|---|---|---|---|---|---|
| 1 | 1 | 400 | 0.71 | nothing | The first measurement becomes the baseline simply because it happened first. |
| 2 | 2 | 400 | 0.74 | nothing | Three points up. If this had followed a prompt edit, the edit would now be considered an improvement and kept. |
| 3 | 3 | 400 | 0.69 | nothing | And back down, five points below the previous run. Still nothing changed. |
| 4 | 4 | 400 | 0.73 | nothing | The spread across four runs is four points, all of it noise. Any explanation offered for any of these moves would be a story about nothing. |
The error on a difference
Now the harder part, which is where most reporting goes wrong. You do not
usually care about one score; you care about the difference between two.
Two independently measured scores each carry their own error, and when you
subtract them the errors do not cancel. They combine: the error on the
difference is the square root of the sum of the squares of the two errors. For
two equal-sized sets that is about 1.4 times the error on either one.
So at 400 cases, with an error of 2.3 points on each score, the error on the
difference is about 3.3 points. A three-point difference is therefore
indistinguishable from zero, and to call it a result you want the difference to
be roughly twice its error.
The lesson stops here
3 more paragraphs to go
You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.
The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.
See the planThe contentsThis is the reading half
Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.
The contents