ContentsThe library

Training a Model From Scratch

Nobody Trains to Convergence

Last timeSpikes, Plateaus and Dead Runs

Runs end because the budget ran out, not because the model converged. Deciding what to keep, and on what evidence, is a separate skill from training.

The run ends when the money does

There is a picture of training in which the loss falls, flattens, and the run

stops because there is nothing left to gain. That picture describes small models

on small data. It describes almost nothing at scale.

At scale the loss is still falling when the run stops, and it stops because the

allotted compute has been spent. The curve has no flat part on it. This changes

what "finished" means: it is not a property the model reaches, it is a decision

somebody makes, and the useful question is not when to stop but how to spend a

fixed amount of compute to get the best model out the other end.

That amount is countable in advance.

FIG 1What a run costs
a constant covering the forward pass, the backward pass and the weight gradient, each of which costs roughly the same arithmetic as the forward pass again
parameters times training tokens, which is the only thing you choose
Floating-point operations for a full training run, accurate enough to plan with. Every decision about a run comes down to picking two numbers whose product is fixed.

Because the product is fixed, model size and data volume trade directly against

each other. Doubling the model halves the tokens it can be trained on. That is

the real decision of a training run, and it is made before a single step runs.

The split, and the rule that came out of it

For several years the common practice made models as large as the hardware would

hold and trained them on whatever data was to hand, which turned out to be a

long way from optimal. Measuring the loss across many combinations of size and

token count, at matched compute, showed that the best split is close to equal

growth in both: a budget twice as large should buy a model about 1.4 times

bigger trained on about 1.4 times more data, not a model twice as big.

The lesson stops here

6 more paragraphs to go

You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.

The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.

See the planThe contents

This is the reading half

Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.

The contents

The rest of this course

  1. 01Four Lines, and Genuinely All of It
  2. 02The Four Hops Between Storage and the Stepopening only
  3. 03The Starting Scale, and Why It Decides Everythingopening only
  4. 04One Number, and Why It Cannot Be One Numberopening only
  5. 05What a Bigger Batch Actually Buysopening only
  6. 06The Four Numbers Worth Watchingopening only
  7. 07Three Ways a Run Breaks, and What Each One Looks Likeopening only
  8. 08Nobody Trains to Convergenceyou are here

Read alongside