The Average Moved Half a Point and the Model Cannot Add Any More
Last timeTeaching a Smaller One
Compression damage is concentrated, not spread. An average score is the one measurement guaranteed to miss it, and here is what to measure instead.
Every technique in this course changes the model and every one of them has a
quality cost. The cost reported in the papers, and the cost most teams measure,
is an average over a broad evaluation set. That number is close to useless for
this decision, and the reason is worth understanding because it is a property
of averages rather than a property of compression.
Why a small average drop means nothing
Suppose a model is evaluated on twenty thousand questions and compression costs
it 0.4 points out of a hundred. That sounds like nothing happened. Now consider
how the 0.4 could be distributed.
It could be eighty items each getting slightly worse. Or it could be eighty
items that went from completely right to completely wrong, while the other
nineteen thousand nine hundred and twenty did not change at all. The average
cannot tell these apart, and the measured answer, consistently, across every
study that has looked, is the second one.
| step | what was measured | size of the set | score before | score after | change | what happened |
|---|---|---|---|---|---|---|
| 1 | a broad standard benchmark | 20000 | 71.2 | 70.8 | -0.4 | The number that gets reported, and the number that gets a compression approved. Within the noise of a rerun. |
| 2 | questions needing four or more reasoning | 600 | 54 | 45.9 | -8.1 | Present in the broad set, but as about three per cent of it, so its collapse contributes a tenth of a point to the average. |
| 3 | recall of entities seen under fifty time | 400 | 62.5 | 51.5 | -11 | A rare fact lives in very few weights. Rounding them is the whole of the damage. |
| 4 | answers needing an exact output format | 500 | 88 | 73 | -15 | The worst one, and the one that breaks a product rather than a score, because downstream code parses that format. |
The arithmetic is simple and worth doing once. A capability present in three
per cent of a set, losing eight points, moves the average by 0.24. A capability
present in half a per cent, losing forty points, moves it by 0.2. A broad
evaluation set is an instrument whose sensitivity to a capability is
proportional to how often that capability appears in it, and the capabilities
compression destroys are by their nature the uncommon ones.
The lesson stops here
5 more paragraphs to go
You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.
The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.
See the planThe contentsThis is the reading half
Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.
The contents