Everything Between the Question and the Number
Last timeThe Gap Between the Name and the Questions
One model, one question set, a dozen defensible ways of running the evaluation, and a spread of fifteen points. The conditions are not a footnote; they are part of the measurement.
Two groups report a score for the same model on the same public question set.
One says 68, the other says 75. Neither is lying, neither made an error, and
both numbers are correct. The difference is in everything that sits between the
question and the number, which neither report described.
The size of the effect
Everyone who runs evaluations knows these choices exist. The usual assumption is
that they are small, a fraction of a point here or there, noise around a stable
underlying ability. The uncomfortable part is that they are not small. They are
larger than the differences most reports are written about, and they are larger
than the sampling error discussed earlier in this course, which at least has the
courtesy of being computable.
| model A | model B | model C | |
|---|---|---|---|
| options lettered, instru | 71 | 68 | 74 |
| options numbered, instru | 64 | 70 | 66 |
| options lettered, instru | 59 | 61 | 70 |
| plain list, no instructi | 66 | 72 | 63 |
One at a time
The effects are easiest to see when the settings are changed singly.
The lesson stops here
3 more paragraphs to go
You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.
The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.
See the planThe contentsThis is the reading half
Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.
The contents