ContentsThe library

Making a Model Smaller

The Average Moved Half a Point and the Model Cannot Add Any More

Last timeTeaching a Smaller One

Compression damage is concentrated, not spread. An average score is the one measurement guaranteed to miss it, and here is what to measure instead.

Every technique in this course changes the model and every one of them has a

quality cost. The cost reported in the papers, and the cost most teams measure,

is an average over a broad evaluation set. That number is close to useless for

this decision, and the reason is worth understanding because it is a property

of averages rather than a property of compression.

Why a small average drop means nothing

Suppose a model is evaluated on twenty thousand questions and compression costs

it 0.4 points out of a hundred. That sounds like nothing happened. Now consider

how the 0.4 could be distributed.

It could be eighty items each getting slightly worse. Or it could be eighty

items that went from completely right to completely wrong, while the other

nineteen thousand nine hundred and twenty did not change at all. The average

cannot tell these apart, and the measured answer, consistently, across every

study that has looked, is the second one.

FIG 1The same compressed model, measured four ways
stepwhat was measuredsize of the setscore beforescore afterchangewhat happened
1a broad standard benchmark2000071.270.8-0.4The number that gets reported, and the number that gets a compression approved. Within the noise of a rerun.
2questions needing four or more reasoning6005445.9-8.1Present in the broad set, but as about three per cent of it, so its collapse contributes a tenth of a point to the average.
3recall of entities seen under fifty time40062.551.5-11A rare fact lives in very few weights. Rounding them is the whole of the damage.
4answers needing an exact output format5008873-15The worst one, and the one that breaks a product rather than a score, because downstream code parses that format.
4 steps
One model, one compression, five measurements. The first line is the one in the report. The fourth line is the one that pages somebody at two in the morning. Nothing about the first line is wrong: it is a correct average, and it is an average over a population in which the damaged cases are rare.

The arithmetic is simple and worth doing once. A capability present in three

per cent of a set, losing eight points, moves the average by 0.24. A capability

present in half a per cent, losing forty points, moves it by 0.2. A broad

evaluation set is an instrument whose sensitivity to a capability is

proportional to how often that capability appears in it, and the capabilities

compression destroys are by their nature the uncommon ones.

The lesson stops here

5 more paragraphs to go

You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.

The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.

See the planThe contents

This is the reading half

Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.

The contents

The rest of this course

  1. 01Three Budgets, and Only One of Them Is the Weights
  2. 02Round It, Store the Integer, Multiply It Backopening only
  3. 03Most of the Model Does Not Care, and You Have to Find the Part That Doesopening only
  4. 04Round the Finished Model, or Train One That Expects Itopening only
  5. 05A Zero Still Occupies Its Place in the Rowopening only
  6. 06The Wrong Answers Are Where the Teaching Isopening only
  7. 07The Average Moved Half a Point and the Model Cannot Add Any Moreyou are here
  8. 08The Savings Multiply and So Does the Damageopening only

Read alongside