Adapting a Model You Did Not Train
The Obvious Approach and What It Costs
Last timeWhat You Are Given and What You Are Not
Continuing to train a model on your own data is the direct way to adapt it. It works, it sets the quality ceiling, and it needs several times the model size in memory to do.
The direct way to change a model is to keep training it, on your data instead of
the internet. There is nothing clever about it and it is worth understanding
properly, because everything cheaper is defined by which part of it has been
given up.
It is just training, continued
The procedure is the one that produced the model. Show it an example, compare
what it produced against what you wanted, work out for every weight whether
raising or lowering it would have helped, and move each one a little in that
direction. Repeat.
Two things differ. The data is yours and there is far less of it. And the step
size is much smaller, typically a hundredth of what pretraining used, because the
model is already close to right and the goal is a correction rather than a
construction.
The memory, in detail
The optimiser is the surprise. The common choice keeps two running averages per
weight, one of the gradient and one of its square, and both are the full size of
the model. So before any intermediate values are stored, four copies exist: the
weights, their gradients, and the two averages.
The lesson stops here
6 more paragraphs to go
You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.
The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.
See the planThe contentsThis is the reading half
Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.
The contents