ContentsThe library

Training a Model From Scratch

Four Lines, and Genuinely All of It

Training is one short loop repeated: predict, score, measure the blame, move the weights. Everything hard about training lives outside those four lines.

The loop, written out

Training a model is one loop. Not a metaphor for a loop: an actual while that

runs a few hundred thousand times and does the same four things each time

round.

Take a batch of examples. Run them through the model as it currently stands and

get predictions out. That is the forward pass, and it is the only step that uses

the model the way a user would. Score the predictions against what the answers

should have been, which produces one number, the loss. Work out, for every

single weight in the model, how much that weight contributed to the loss being

as large as it was. That is the gradient, and it is the same size as the model.

Then nudge every weight a small distance in the direction that would have made

the loss smaller.

That is the whole of it. Four steps, repeated. There is no fifth step in which

the model consolidates what it has learnt, and no stage at which anything is

stored anywhere other than in the weights themselves.

FIG 1One turn of the loop
The arrow back to the top is the entire structure. Nothing accumulates anywhere except in the weights.

The update step is worth writing down, because the simplest version of it is

one line and every more elaborate optimiser is a variation on it.

FIG 2The update
every weight in the model, as it stands at step t
the learning rate: how far to move. One number, and the one that decides most
the gradient, which is how much the loss would rise per unit increase in each weight
The minus sign is the whole idea: move against the direction that would make the score worse.

The loss is the only opinion in the room

Look at the loop again and ask where any preference lives. The forward pass has

none: it computes what the weights say it computes. The gradient has none: it

reports a fact about the arithmetic. The update has none: it moves against the

gradient by a fixed distance.

The loss is the only place in the loop where anything is called good or bad. So

the loss is not one component among several. It is the complete specification of

what the model will become, and everything else is machinery for satisfying it.

This has a consequence people learn the hard way. A model trained to predict the

next token will become extremely good at predicting the next token, including

the next token of text that is wrong, dated or hostile, because the loss never

mentioned truth. A model scored on whether a human rater liked the answer will

become good at producing answers raters like, which overlaps with being correct

but is not the same set. Neither of these is a failure of training. Both runs

minimised exactly what they were told to minimise.

FIG 3The loop with nothing hidden
python
for step in range(steps):
    batch = next(data)              # one batch, already on the device
    out   = model(batch.inputs)     # forward
    loss  = objective(out, batch.targets)
    loss.backward()                 # gradient for every weight
    optimiser.step()                # move
    optimiser.zero_grad()           # forget this batch's gradient
Real training code is longer than this, but not because the loop grew: because of logging, checkpointing and the data pipeline around it.

The last line catches people. The gradient accumulates by default, so forgetting

to clear it means step two is corrected using the sum of two batches, step three

using three, and the effective step size grows without anybody changing the

learning rate. The loss curve wanders upward and looks like a learning-rate

problem.

Step, batch, epoch

Three words count three different things and reports use them interchangeably.

A step is one pass through the loop: one update to the weights. A batch

is the set of examples used in one step. An epoch is one complete pass

through the training data, which is however many steps it takes to see all of

it once.

So if the data holds a million examples and the batch holds a thousand, an epoch

is a thousand steps. Double the batch and an epoch is five hundred steps, while

the amount of data seen is identical. This is why "trained for three epochs" and

"trained for thirty thousand steps" are not comparable claims unless the batch

size is also given, and why a report that gives only one of the three is

unreadable.

FIG 4The same run counted three ways
stepstepexamples seenepochwhat happened
1110000.001One update. The model has changed, slightly, on the evidence of a thousand examples.
25005000000.5Half a lap. Half the data has never been seen at all.
3100010000001One epoch, on a million-example set with a batch of a thousand.
4300030000003Three epochs. Every example has now contributed to three separate updates.
4 steps
A million examples, a batch of a thousand. Reports quote whichever of these three columns flatters them.

For very large models the epoch has quietly stopped being a useful unit at all,

because the data is larger than the run: a model may be trained on a trillion

tokens and never complete one lap. Steps and tokens are what get quoted, and

epochs appear only in work on smaller sets where repeating the data is normal.

The loop is almost never the fault

Here is the practical reason to be this precise about something so short. When a

run is not learning, the loop is the last place to look, because the loop is

four calls into libraries that are exercised by millions of other runs every

day.

Three other places account for nearly all of it. The data: a pipeline that is

feeding the same batch repeatedly, or inputs and targets that have come apart by

one position, or a field that is empty in the stored format and silently becomes

zeros. The step size: too large and the loss rises or goes to nothing at all,

too small and the curve is a flat line that looks like a bug. And the plumbing

around the loop: a gradient that was never cleared, a model left in evaluation

mode, a loss averaged over the wrong axis.

Each of the remaining lessons is one of those places. The loop does not come up

again, because there is nothing more to say about it.

Recap

  • Training is a loop of four steps, and the loop itself is short enough to write out in full on one page.
  • The loss function is the entire specification of what the model will become, because nothing else in the loop has an opinion about what is good.
  • Almost every training failure is in the data feeding the loop or in the numbers configuring it, not in the loop.

This is the reading half

Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.

The contents

NextGetting the Data to the Chip →

The rest of this course

  1. 01Four Lines, and Genuinely All of Ityou are here
  2. 02The Four Hops Between Storage and the Stepopening only
  3. 03The Starting Scale, and Why It Decides Everythingopening only
  4. 04One Number, and Why It Cannot Be One Numberopening only
  5. 05What a Bigger Batch Actually Buysopening only
  6. 06The Four Numbers Worth Watchingopening only
  7. 07Three Ways a Run Breaks, and What Each One Looks Likeopening only
  8. 08Nobody Trains to Convergenceopening only

Read alongside