ContentsThe library

Making a Model Smaller

The Wrong Answers Are Where the Teaching Is

Last timeTaking Weights Out Entirely

A small model trained on a large one's full output distribution learns more than the same model trained on the original labels. Here is the objective and why it works.

Every technique so far has taken a large model and described it more cheaply.

This one builds a different, smaller model and teaches it to behave like the

large one. The result is genuinely small: fewer layers, narrower, less state

per request, faster arithmetic. All three budgets from the first lesson fall at

once.

The interesting part is not the smaller architecture, which anybody can

specify. It is that training the small model on the large one's outputs works

considerably better than training the same small model on the original data,

and the reason is worth understanding because it tells you what to match.

What a label leaves out

Take one position in a sentence: "the capital of Australia is ___". The

training data says the answer is Canberra. That is the whole of what the label

communicates: one token is right and the other fifty thousand are wrong, with

no distinction among them.

The teacher, asked the same question, produces a probability for every token in

its vocabulary.

FIG 1A teacher's answer to one question
The label says only that the first slice is correct. The teacher says that two wrong answers were very nearly right, that a handful of others were in the right category, and that the remaining fifty thousand tokens were not under consideration at all. That ranking is a summary of what the teacher learned about the world, and it is delivered for every token of every input.

That ranking is the thing being transferred. It says Sydney and Melbourne are

the kind of thing that goes here; it says the answer is a city; it says

punctuation and verbs are not candidates. The label says none of it, and a

small model learning from labels alone has to rediscover all of it from scratch

with far less capacity than the teacher had.

The lesson stops here

4 more paragraphs to go

You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.

The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.

See the planThe contents

This is the reading half

Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.

The contents

The rest of this course

  1. 01Three Budgets, and Only One of Them Is the Weights
  2. 02Round It, Store the Integer, Multiply It Backopening only
  3. 03Most of the Model Does Not Care, and You Have to Find the Part That Doesopening only
  4. 04Round the Finished Model, or Train One That Expects Itopening only
  5. 05A Zero Still Occupies Its Place in the Rowopening only
  6. 06The Wrong Answers Are Where the Teaching Isyou are here
  7. 07The Average Moved Half a Point and the Model Cannot Add Any Moreopening only
  8. 08The Savings Multiply and So Does the Damageopening only

Read alongside