ContentsThe library

Making a Model Smaller

Most of the Model Does Not Care, and You Have to Find the Part That Does

Last timeKeeping Fewer Bits

Rounding damage is not spread evenly. A few per cent of a model carries most of the sensitivity, and keeping that part precise costs almost nothing.

The last lesson produced a procedure that can be applied to any weight in a

model. The obvious thing to do with it is apply it to all of them at the same

setting, and that is what the simplest tools do. It is also reliably the wrong

thing, because the damage from rounding is not spread evenly across a model.

The uneven part is sharp rather than gentle. Two layers of identical shape, a

few positions apart in the same network, can differ by more than tenfold in how

much the model degrades when you round them and leave everything else alone. A

single setting applied to both over-rounds one and wastes memory on the other.

FIG 1Rounding one part at a time, everything else left exact
Share of all weights, peQuality cost alone, poinQuality cost per per cenKept precise in practice
Embedding table12.000.400.030.00
First two layers5.003.100.621.00
Middle layers, weights58.001.100.020.00
Middle layers, attention16.002.400.151.00
Final layer3.002.800.931.00
Output table6.004.200.701.00
A seven billion parameter model, each part rounded to four bits on its own. The third column is the one that decides anything: cost per per cent of the model. The two marked rows are thirty times more sensitive than the middle layers, and together they are nine per cent of the weights. Keeping just those two precise is almost free and recovers most of the loss.

What the uneven damage looks like

Read the first and third columns together. The middle layers are the bulk of

the model and are nearly indifferent: fifty-eight per cent of the weights for

about a point of damage. The output table is six per cent of the weights for

four points. If you have a memory budget and are deciding where to spend

precision, that ratio is the whole decision, and it is not guessable from the

shapes of the layers.

Where it reliably hurts

Four places come up again and again, and knowing them tells you where to start

measuring rather than what to conclude.

The first and last layers sit at the boundary with the outside world. The first

converts tokens into the representation everything else works in, and an error

there propagates through every subsequent layer. The last converts back into a

score for every token in the vocabulary, where small differences decide which

word comes out. Errors in the middle have layers after them that can partly

absorb the damage; errors at the ends do not.

The embedding and output tables are large and look like obvious targets, and

they behave differently from each other. The embedding table tolerates rounding

well, because a small shift in a token's representation is something the model

is robust to anyway. The output table does not, because it is compared against

every candidate and the comparison is close.

The attention projections are more sensitive than the larger blocks that follow

them, which is counterintuitive given their size and is reproducible across

models.

And then the one people miss.

The lesson stops here

5 more paragraphs to go

You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.

The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.

See the planThe contents

This is the reading half

Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.

The contents

The rest of this course

  1. 01Three Budgets, and Only One of Them Is the Weights
  2. 02Round It, Store the Integer, Multiply It Backopening only
  3. 03Most of the Model Does Not Care, and You Have to Find the Part That Doesyou are here
  4. 04Round the Finished Model, or Train One That Expects Itopening only
  5. 05A Zero Still Occupies Its Place in the Rowopening only
  6. 06The Wrong Answers Are Where the Teaching Isopening only
  7. 07The Average Moved Half a Point and the Model Cannot Add Any Moreopening only
  8. 08The Savings Multiply and So Does the Damageopening only

Read alongside