Four Lines, and Genuinely All of It
Training is one short loop repeated: predict, score, measure the blame, move the weights. Everything hard about training lives outside those four lines.
The loop, written out
Training a model is one loop. Not a metaphor for a loop: an actual while that
runs a few hundred thousand times and does the same four things each time
round.
Take a batch of examples. Run them through the model as it currently stands and
get predictions out. That is the forward pass, and it is the only step that uses
the model the way a user would. Score the predictions against what the answers
should have been, which produces one number, the loss. Work out, for every
single weight in the model, how much that weight contributed to the loss being
as large as it was. That is the gradient, and it is the same size as the model.
Then nudge every weight a small distance in the direction that would have made
the loss smaller.
That is the whole of it. Four steps, repeated. There is no fifth step in which
the model consolidates what it has learnt, and no stage at which anything is
stored anywhere other than in the weights themselves.
The update step is worth writing down, because the simplest version of it is
one line and every more elaborate optimiser is a variation on it.
- every weight in the model, as it stands at step t
- the learning rate: how far to move. One number, and the one that decides most
- the gradient, which is how much the loss would rise per unit increase in each weight
The loss is the only opinion in the room
Look at the loop again and ask where any preference lives. The forward pass has
none: it computes what the weights say it computes. The gradient has none: it
reports a fact about the arithmetic. The update has none: it moves against the
gradient by a fixed distance.
The loss is the only place in the loop where anything is called good or bad. So
the loss is not one component among several. It is the complete specification of
what the model will become, and everything else is machinery for satisfying it.
This has a consequence people learn the hard way. A model trained to predict the
next token will become extremely good at predicting the next token, including
the next token of text that is wrong, dated or hostile, because the loss never
mentioned truth. A model scored on whether a human rater liked the answer will
become good at producing answers raters like, which overlaps with being correct
but is not the same set. Neither of these is a failure of training. Both runs
minimised exactly what they were told to minimise.
for step in range(steps):
batch = next(data) # one batch, already on the device
out = model(batch.inputs) # forward
loss = objective(out, batch.targets)
loss.backward() # gradient for every weight
optimiser.step() # move
optimiser.zero_grad() # forget this batch's gradientThe last line catches people. The gradient accumulates by default, so forgetting
to clear it means step two is corrected using the sum of two batches, step three
using three, and the effective step size grows without anybody changing the
learning rate. The loss curve wanders upward and looks like a learning-rate
problem.
Step, batch, epoch
Three words count three different things and reports use them interchangeably.
A step is one pass through the loop: one update to the weights. A batch
is the set of examples used in one step. An epoch is one complete pass
through the training data, which is however many steps it takes to see all of
it once.
So if the data holds a million examples and the batch holds a thousand, an epoch
is a thousand steps. Double the batch and an epoch is five hundred steps, while
the amount of data seen is identical. This is why "trained for three epochs" and
"trained for thirty thousand steps" are not comparable claims unless the batch
size is also given, and why a report that gives only one of the three is
unreadable.
| step | step | examples seen | epoch | what happened |
|---|---|---|---|---|
| 1 | 1 | 1000 | 0.001 | One update. The model has changed, slightly, on the evidence of a thousand examples. |
| 2 | 500 | 500000 | 0.5 | Half a lap. Half the data has never been seen at all. |
| 3 | 1000 | 1000000 | 1 | One epoch, on a million-example set with a batch of a thousand. |
| 4 | 3000 | 3000000 | 3 | Three epochs. Every example has now contributed to three separate updates. |
For very large models the epoch has quietly stopped being a useful unit at all,
because the data is larger than the run: a model may be trained on a trillion
tokens and never complete one lap. Steps and tokens are what get quoted, and
epochs appear only in work on smaller sets where repeating the data is normal.
The loop is almost never the fault
Here is the practical reason to be this precise about something so short. When a
run is not learning, the loop is the last place to look, because the loop is
four calls into libraries that are exercised by millions of other runs every
day.
Three other places account for nearly all of it. The data: a pipeline that is
feeding the same batch repeatedly, or inputs and targets that have come apart by
one position, or a field that is empty in the stored format and silently becomes
zeros. The step size: too large and the loss rises or goes to nothing at all,
too small and the curve is a flat line that looks like a bug. And the plumbing
around the loop: a gradient that was never cleared, a model left in evaluation
mode, a loss averaged over the wrong axis.
Each of the remaining lessons is one of those places. The loop does not come up
again, because there is nothing more to say about it.
Recap
- Training is a loop of four steps, and the loop itself is short enough to write out in full on one page.
- The loss function is the entire specification of what the model will become, because nothing else in the loop has an opinion about what is good.
- Almost every training failure is in the data feeding the loop or in the numbers configuring it, not in the loop.
This is the reading half
Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.
The contents