Adapting a Model You Did Not Train
Almost All of the Work Is Already Done
A pretrained model already knows how language works and what things are called. What it lacks for your task is a small correction, and the whole subject is about paying only for that.
Nobody in the position of wanting a model for a particular job starts by
training one. They start with a model somebody else trained, and the interesting
question is what that gives them and what it leaves to do.
What the expensive part bought
Pretraining a model consists of showing it an enormous amount of text and asking
it to predict what comes next. It is a strange objective for such a general
result, but it produces a model that knows how sentences are shaped, what words
refer to, which facts commonly appear together, and how an argument tends to
proceed.
None of that is specific to any one task, and all of it would have to be learned
again by anyone starting from random weights. It is also where essentially all
the money went.
How far it has to move
The useful surprise is that the change a task needs is not just cheaper than
pretraining, it is small in an absolute sense that can be measured.
Take a model and allow it to change only within a restricted set of directions
of a chosen size, rather than freely in every direction its weights offer. Shrink
that set until quality starts to suffer, and the size at which it does is a
measurement of how big the task's correction really is.
| millions of weights | directions needed for ni | share of full quality re | |
|---|---|---|---|
| a small model | 125.0 | 1861.0 | 0.9 |
| a medium model | 355.0 | 1326.0 | 0.9 |
| a larger model | 774.0 | 1021.0 | 0.9 |
That result is the licence for everything later in this course. If the
correction genuinely needs only a few thousand numbers, then a method that stores
a few thousand numbers has lost nothing, and methods that store hundreds of
millions were paying for capacity the task could not use.
- a weight matrix of the pretrained model, which you were given and did not pay for
- the correction this task requires, the same shape as W but far simpler
- the adapted weight matrix, which is what actually runs
It is also worth being clear that the correction is a correction and not a
replacement. The adapted model is the original model plus something, and the
original is still in there doing most of the work. That is why adaptation is
fast, why it needs little data, and also why it can quietly damage abilities
nobody was trying to change.
Why starting elsewhere works at all
It is worth asking why a model trained to predict ordinary text should be a good
starting point for classifying support tickets or writing in a house style. The
answer is that most of what such a model computes is not about the text it was
trained on.
Three different things people want
When somebody says a model is not good enough for their purpose they usually
mean one of three quite different things, and the distinction decides what is
worth doing.
The first is that the model does not know something. It has never seen your
internal documentation and cannot answer questions about it. The second is that
the model does not behave as wanted: the output is the wrong shape, the register
is wrong, it refuses things it should not or explains when it should answer. The
third is that the model's judgement is poor on your particular problem, grading
work too leniently or classifying ambiguous cases wrongly.
Only the second and third are things training on your data does well. Teaching
behaviour takes surprisingly few examples, because the model already has every
ability involved and only needs to be told which to use. Judgement takes more
and is genuinely learnable. Facts are the awkward case: training can push a few
in, but it is unreliable, it damages other things, and retrieving the document
at the time of the question is almost always better.
What to hold on to
You are given a general model that cost a fortune and knows almost everything
except your situation. The correction your task needs is small enough to write
down in a few thousand numbers. What remains is to decide where those numbers
live, what they are allowed to change, and whether your problem is one that
training fixes at all.
Recap
- Pretraining buys a general representation of language that would cost millions to produce and that your task does not need to pay for again, so adaptation starts from somewhere very close to where it needs to end.
- The correction a specific task requires is small in a precise and measurable sense: it can be confined to a few thousand directions of change and still reach most of the quality that changing everything reaches.
- What you actually want fixed matters, because missing knowledge, wrong format and poor judgement are three different problems and only one of them is reliably solved by training on more of your own data.
This is the reading half
Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.
The contents