ContentsThe library

Training a Model From Scratch

One Number, and Why It Cannot Be One Number

Last timeWhere the Weights Start

The learning rate changes the outcome of a run more than any architectural choice, and the right value at step one is the wrong value at step fifty thousand.

The narrowest dial on the machine

Of all the numbers configuring a training run, the learning rate is the one that

decides most. Change the number of layers by twenty per cent and the run is

slightly better or slightly worse. Change the learning rate by a factor of ten

and the run either works or produces nothing at all.

The band of workable values is narrow, and it is narrow in both directions.

FIG 1Where a learning rate can live
trains, but will not finish in your budgetthe band that worksdiverges, usually within 500 steps-7-6-2-1-5too slow to finish-4typical for a large model-3upper edgelearning rate, as a power of ten
Roughly two orders of magnitude of workable values out of six, and only the middle of that band is comfortable.

The two failures look nothing alike, which is the one piece of good news. Too

large and the loss rises, oscillates violently, or becomes a value that is not a

number, almost always within the first few hundred steps. Too small and the loss

curve is a nearly straight line with a slight downward tilt, which is easy to

mistake for a model that cannot learn the task.

That second failure is the expensive one, because it does not look like a

failure. It looks like a result. The test is cheap: multiply by ten, run three

hundred steps, and see whether the curve bends.

Why the start is different

Set aside the schedule and ask what is unusual about step one.

The weights are random, so the model's predictions are essentially arbitrary and

the loss is at its maximum. Large loss means large gradients, so the very first

updates are the largest moves the run will ever make. And the optimiser, if it

is one of the adaptive ones, maintains running averages of the gradient and of

its square in order to scale each weight's step. At step one those averages are

built from one observation. At step ten, from ten. Their estimates are at their

least reliable exactly when the moves are at their largest.

Each of those on its own argues for taking the first steps gently. Together they

are the reason warmup exists: hold the learning rate near zero and ramp it

linearly up to its intended value over the first few hundred to few thousand

steps.

The lesson stops here

5 more paragraphs to go

You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.

The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.

See the planThe contents

This is the reading half

Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.

The contents

The rest of this course

  1. 01Four Lines, and Genuinely All of It
  2. 02The Four Hops Between Storage and the Stepopening only
  3. 03The Starting Scale, and Why It Decides Everythingopening only
  4. 04One Number, and Why It Cannot Be One Numberyou are here
  5. 05What a Bigger Batch Actually Buysopening only
  6. 06The Four Numbers Worth Watchingopening only
  7. 07Three Ways a Run Breaks, and What Each One Looks Likeopening only
  8. 08Nobody Trains to Convergenceopening only

Read alongside