ContentsThe library

Adapting a Model You Did Not Train

Almost All of the Work Is Already Done

A pretrained model already knows how language works and what things are called. What it lacks for your task is a small correction, and the whole subject is about paying only for that.

Nobody in the position of wanting a model for a particular job starts by

training one. They start with a model somebody else trained, and the interesting

question is what that gives them and what it leaves to do.

What the expensive part bought

Pretraining a model consists of showing it an enormous amount of text and asking

it to predict what comes next. It is a strange objective for such a general

result, but it produces a model that knows how sentences are shaped, what words

refer to, which facts commonly appear together, and how an argument tends to

proceed.

None of that is specific to any one task, and all of it would have to be learned

again by anyone starting from random weights. It is also where essentially all

the money went.

FIG 1Where the compute goes, across a model lifetime
The second slice is barely visible, and that is the point. Adaptation is a rounding error against the cost of the thing being adapted, which is why the question is never whether to start from a pretrained model but only how to change it.

How far it has to move

The useful surprise is that the change a task needs is not just cheaper than

pretraining, it is small in an absolute sense that can be measured.

Take a model and allow it to change only within a restricted set of directions

of a chosen size, rather than freely in every direction its weights offer. Shrink

that set until quality starts to suffer, and the size at which it does is a

measurement of how big the task's correction really is.

FIG 2Directions of change needed, against the size of the model
millions of weightsdirections needed for nishare of full quality re
a small model125.01861.00.9
a medium model355.01326.00.9
a larger model774.01021.00.9
Around a thousand numbers are enough to specify what the task requires, inside a model holding hundreds of millions. The count also falls as the pretrained model gets better, which is the opposite of the intuition that a bigger model needs a bigger correction.

That result is the licence for everything later in this course. If the

correction genuinely needs only a few thousand numbers, then a method that stores

a few thousand numbers has lost nothing, and methods that store hundreds of

millions were paying for capacity the task could not use.

FIG 3What adaptation is, before any method is chosen
a weight matrix of the pretrained model, which you were given and did not pay for
the correction this task requires, the same shape as W but far simpler
the adapted weight matrix, which is what actually runs
Every method in this course is a different answer to one question: what is allowed to go in the middle term, and where is it stored. Changing everything lets it be any matrix at all. The cheap methods restrict its form, and the measurement above says that restriction costs much less than it sounds like it should.

It is also worth being clear that the correction is a correction and not a

replacement. The adapted model is the original model plus something, and the

original is still in there doing most of the work. That is why adaptation is

fast, why it needs little data, and also why it can quietly damage abilities

nobody was trying to change.

Why starting elsewhere works at all

It is worth asking why a model trained to predict ordinary text should be a good

starting point for classifying support tickets or writing in a house style. The

answer is that most of what such a model computes is not about the text it was

trained on.

FIG 4What moves and what stays when a model is adapted
The division is not sharp and the exact shape depends on the task, but the direction is reliable. Adaptation is mostly a change to how existing understanding is used, not to the understanding itself.

Three different things people want

When somebody says a model is not good enough for their purpose they usually

mean one of three quite different things, and the distinction decides what is

worth doing.

The first is that the model does not know something. It has never seen your

internal documentation and cannot answer questions about it. The second is that

the model does not behave as wanted: the output is the wrong shape, the register

is wrong, it refuses things it should not or explains when it should answer. The

third is that the model's judgement is poor on your particular problem, grading

work too leniently or classifying ambiguous cases wrongly.

FIG 5Quality against the amount of your own labelled data
0.000.300.600.901.2050.01287.52525.03762.55000.0labelled examples available
starting from a pretrained modelstarting from random weights
The pretrained curve starts high and rises quickly, which is why a few hundred examples is often enough to teach a behaviour. Both curves flatten, and neither of them is a way to install facts the model has never seen, which is the first want and the one training serves worst.

Only the second and third are things training on your data does well. Teaching

behaviour takes surprisingly few examples, because the model already has every

ability involved and only needs to be told which to use. Judgement takes more

and is genuinely learnable. Facts are the awkward case: training can push a few

in, but it is unreliable, it damages other things, and retrieving the document

at the time of the question is almost always better.

What to hold on to

You are given a general model that cost a fortune and knows almost everything

except your situation. The correction your task needs is small enough to write

down in a few thousand numbers. What remains is to decide where those numbers

live, what they are allowed to change, and whether your problem is one that

training fixes at all.

Recap

  • Pretraining buys a general representation of language that would cost millions to produce and that your task does not need to pay for again, so adaptation starts from somewhere very close to where it needs to end.
  • The correction a specific task requires is small in a precise and measurable sense: it can be confined to a few thousand directions of change and still reach most of the quality that changing everything reaches.
  • What you actually want fixed matters, because missing knowledge, wrong format and poor judgement are three different problems and only one of them is reliably solved by training on more of your own data.

This is the reading half

Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.

The contents

NextChanging Every Weight →

The rest of this course

  1. 01Almost All of the Work Is Already Doneyou are here
  2. 02The Obvious Approach and What It Costsopening only
  3. 03The Abilities Nobody Was Trying to Changeopening only
  4. 04Writing the Correction as Two Thin Sheetsopening only
  5. 05Spending a Fixed Budget in the Right Placesopening only
  6. 06Four Other Places to Put the Changeopening only
  7. 07Keeping the Correction Separate on Purposeopening only
  8. 08Deciding What Actually Needs Changingopening only

Read alongside