Most of the Model Does Not Care, and You Have to Find the Part That Does
Last timeKeeping Fewer Bits
Rounding damage is not spread evenly. A few per cent of a model carries most of the sensitivity, and keeping that part precise costs almost nothing.
The last lesson produced a procedure that can be applied to any weight in a
model. The obvious thing to do with it is apply it to all of them at the same
setting, and that is what the simplest tools do. It is also reliably the wrong
thing, because the damage from rounding is not spread evenly across a model.
The uneven part is sharp rather than gentle. Two layers of identical shape, a
few positions apart in the same network, can differ by more than tenfold in how
much the model degrades when you round them and leave everything else alone. A
single setting applied to both over-rounds one and wastes memory on the other.
| Share of all weights, pe | Quality cost alone, poin | Quality cost per per cen | Kept precise in practice | |
|---|---|---|---|---|
| Embedding table | 12.00 | 0.40 | 0.03 | 0.00 |
| First two layers | 5.00 | 3.10 | 0.62 | 1.00 |
| Middle layers, weights | 58.00 | 1.10 | 0.02 | 0.00 |
| Middle layers, attention | 16.00 | 2.40 | 0.15 | 1.00 |
| Final layer | 3.00 | 2.80 | 0.93 | 1.00 |
| Output table | 6.00 | 4.20 | 0.70 | 1.00 |
What the uneven damage looks like
Read the first and third columns together. The middle layers are the bulk of
the model and are nearly indifferent: fifty-eight per cent of the weights for
about a point of damage. The output table is six per cent of the weights for
four points. If you have a memory budget and are deciding where to spend
precision, that ratio is the whole decision, and it is not guessable from the
shapes of the layers.
Where it reliably hurts
Four places come up again and again, and knowing them tells you where to start
measuring rather than what to conclude.
The first and last layers sit at the boundary with the outside world. The first
converts tokens into the representation everything else works in, and an error
there propagates through every subsequent layer. The last converts back into a
score for every token in the vocabulary, where small differences decide which
word comes out. Errors in the middle have layers after them that can partly
absorb the damage; errors at the ends do not.
The embedding and output tables are large and look like obvious targets, and
they behave differently from each other. The embedding table tolerates rounding
well, because a small shift in a token's representation is something the model
is robust to anyway. The output table does not, because it is compared against
every candidate and the comparison is close.
The attention projections are more sensitive than the larger blocks that follow
them, which is counterintuitive given their size and is reproducible across
models.
And then the one people miss.
The lesson stops here
5 more paragraphs to go
You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.
The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.
See the planThe contentsThis is the reading half
Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.
The contents