How Far the Number Could Have Landed From Where It Did
Last timeA Score Is a Sample
A score measured on a few hundred questions carries several points of wobble that come from nothing but which questions were drawn. Here is how to work out how many.
A model scores 84 on a set of a hundred questions. Somebody writes a hundred
more questions in exactly the same way, from the same source, with the same
care. The same model scores 79. Nothing has changed except which questions were
drawn, and the number has moved five points.
Where the movement comes from
This is not a flaw in the benchmark and not noise in the model. It is what
happens whenever you measure a proportion on a sample. Each question is a coin
that lands right or wrong, the score is the share that landed right, and the
share from one handful of coins differs from the share from the next.
- the standard error: the typical distance between the measured score and the true one
- the score you measured, as a fraction of one
- how many questions were asked
Put the numbers in. A score of 0.84 on a hundred questions gives a standard
error of about 0.037, which is 3.7 points. That is the typical distance. The
range people mean when they draw an error bar is about twice that in each
direction, so the honest statement is not eighty four but somewhere between
roughly seventy seven and ninety one.
The lesson stops here
6 more paragraphs to go
You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.
The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.
See the planThe contentsThis is the reading half
Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.
The contents