ContentsThe library

Keeping The Numbers In Range

What A Training Curve Is Telling You

Last timeThe Same Trouble Backwards

Everything in this course arrives in practice as one line on a chart going the wrong way, so the last skill is reading that line and knowing which of the remedies to reach for.

Seven lessons of mechanism arrive in practice as one line on a chart that is not

doing what it should. This lesson is about that line: what it can tell you, what

it cannot, and what to measure next.

Start with the honest limitation. The loss curve is the cheapest instrument you

have, because it is recorded anyway, and it is also the least specific. A curve

that falls and then flattens at a disappointing value is consistent with a

vanishing gradient, a step size far too small, a model too small for the task, and

a mistake in how the data is labelled. Those four have nothing in common and the

curve does not distinguish them.

FIG 1A run with a spike in it
2.004.507.009.5012.001.010.820.530.340.0thousands of steps
the loss as recorded
The shape to recognise. A normal descent, a sharp rise over a few hundred steps, and a return to roughly where it was. On a run of this size the spike is obvious. On a chart averaged over each pass through the data it would be a small bump or nothing at all, which is the first argument for recording the loss per batch rather than per epoch: a fault that lasts six steps has to be looked for at the resolution of steps.

What a spike is and what is usually done about it

A spike is a sudden large rise in a run that was descending normally. An enormous

gradient arrived, the step moved the parameters somewhere much worse, and the run

has to climb back out.

The published accounts of large training runs are unusually consistent here. One

reports roughly twenty spikes in a single run and describes the response: go back

to a saved state from a few hundred steps before the spike, skip the batches that

were being processed when it happened, and continue. Interestingly, the same

batches usually pass without incident when the run reaches them from a slightly

different state, which says the batch was not corrupt. It was an unlucky

interaction between a particular batch and a particular set of parameters.

FIG 2A spike, and the recovery
stepsteplossoverall gradient lengthwhat was donewhat happened
1214002.610.42nothing, a normal stepThe run is descending at the expected rate and the gradient length is typical for this stage.
2214182.6418clipped to the thresholdA large gradient arrives. Clipping bounds the step, and on most occasions this is the end of the story.
3214193.956.1nothingThe loss has risen sharply anyway. The clipped step was still enough to move the parameters into a much worse region.
4214603.42.2nothingThe run is climbing back out on its own, slowly, and will take thousands of steps to recover the ground it lost.
5214002.610.42restarted from the saved state, batches The intervention: reload the state from before the spike, skip the batches around step 21418, and continue from there.
5 steps
Note the second row. Clipping did fire and the spike happened anyway, which is worth knowing in advance: clipping bounds a step without making a bounded step safe. The recovery in the fourth row is real but slow, and the last row is cheaper than waiting for it.

Three numbers worth recording

The loss alone leaves you guessing. Three further measurements, all cheap, turn a

vague curve into a named fault.

The length of the gradient at each step, which shows an explosion directly and also

shows how often clipping is firing. If it fires on most steps the threshold is

acting as a cap on the step size rather than as a safety measure.

The size of each update relative to the size of the parameter it changes. This is

more informative than the gradient on its own, because the optimiser rescales

gradients and a gradient of a given size says little about how far anything

actually moved. It is also the quantity that has been found to predict instability

at larger scale when measured in small runs.

The spread of the values at several depths, which shows the forward problem from

the second lesson and, more usefully, shows where in the stack it starts.

The lesson stops here

2 more paragraphs to go

You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.

The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.

See the planThe contents

This is the reading half

Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.

The contents

The rest of this course

  1. 01The Numbers A Machine Cannot Hold
  2. 02The Compounding Nobody Budgets Foropening only
  3. 03The Decision Made Before Anything Runsopening only
  4. 04Putting The Numbers Back Where They Belongopening only
  5. 05Two Decisions That Look Like Detailsopening only
  6. 06The Road That Goes Roundopening only
  7. 07The Same Trouble, Running Backwardsopening only
  8. 08What A Training Curve Is Telling Youyou are here

Read alongside