Nobody Trains to Convergence
Last timeSpikes, Plateaus and Dead Runs
Runs end because the budget ran out, not because the model converged. Deciding what to keep, and on what evidence, is a separate skill from training.
The run ends when the money does
There is a picture of training in which the loss falls, flattens, and the run
stops because there is nothing left to gain. That picture describes small models
on small data. It describes almost nothing at scale.
At scale the loss is still falling when the run stops, and it stops because the
allotted compute has been spent. The curve has no flat part on it. This changes
what "finished" means: it is not a property the model reaches, it is a decision
somebody makes, and the useful question is not when to stop but how to spend a
fixed amount of compute to get the best model out the other end.
That amount is countable in advance.
- a constant covering the forward pass, the backward pass and the weight gradient, each of which costs roughly the same arithmetic as the forward pass again
- parameters times training tokens, which is the only thing you choose
Because the product is fixed, model size and data volume trade directly against
each other. Doubling the model halves the tokens it can be trained on. That is
the real decision of a training run, and it is made before a single step runs.
The split, and the rule that came out of it
For several years the common practice made models as large as the hardware would
hold and trained them on whatever data was to hand, which turned out to be a
long way from optimal. Measuring the loss across many combinations of size and
token count, at matched compute, showed that the best split is close to equal
growth in both: a budget twice as large should buy a model about 1.4 times
bigger trained on about 1.4 times more data, not a model twice as big.
The lesson stops here
6 more paragraphs to go
You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.
The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.
See the planThe contentsThis is the reading half
Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.
The contents