Adapting a Model You Did Not Train
Writing the Correction as Two Thin Sheets
Last timeWhat Gets Lost On the Way
The correction adaptation needs is far simpler than the model it corrects, so it can be written as a product of two thin matrices and trained with a thousandth of the parameters.
The previous two lessons set up a problem and most of the answer. The
difference between the pretrained model and the adapted one is a correction the
same shape as the weights. The question is whether that correction really needs
to be as complicated as the thing it corrects.
It does not
Pretraining had to learn language, reasoning, and a vast amount of world
knowledge, and that genuinely uses the full capacity of the weights. Adaptation
asks for something much smaller: answer in this format, adopt this tone, use
this vocabulary, classify into these six categories.
The correction that expresses such a target repeats itself heavily. Work that
measures how many independent directions an adaptation actually moves in finds
the number is startlingly small, often a few hundred out of hundreds of
millions, and that it shrinks as the pretrained model gets better. That is a
measured fact about tasks, not a modelling assumption anybody had to defend.
The lesson stops here
7 more paragraphs to go
You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.
The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.
See the planThe contentsThis is the reading half
Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.
The contents