ContentsThe library

Training a Model From Scratch

What a Bigger Batch Actually Buys

Last timeHow Fast to Move, and When

Batch size sets how noisy each gradient is, and the noise falls only with the square root, which is why doubling a batch buys far less than it costs.

The gradient is an average

The gradient used in a step is computed on a batch, not on the data. Each

example in the batch produces its own gradient, and the batch gradient is their

average. That single observation is enough to settle how batch size behaves,

because an average of samples is a thing statistics has a complete account of.

The error in an average falls as the square root of the number of samples.

FIG 1The error in a batch gradient
how far the batch gradient typically sits from the true gradient of the whole data
how much per-example gradients disagree with each other, a property of the data and the model, not of the batch
the batch size
The root is the whole story: four times the examples per step for half the noise.

Say that out loud as a price. Going from a batch of 256 to a batch of 1024 costs

four times the arithmetic per step and buys a gradient that is twice as

accurate. Going from 1024 to 4096 costs four times again for another halving.

The returns do not merely diminish, they diminish on a known schedule.

FIG 2What each doubling buys
0.000.300.600.901.2016.0268.0520.0772.01024.0batch size
error
Most of the available noise reduction happens in the first few doublings, and the curve is nearly flat by the right-hand edge.

Why the step size has to move with it

Here is the mistake almost everybody makes once. Take a working run, multiply

the batch size by four to use the hardware better, change nothing else, and

observe that the result is worse.

Count the steps. The data is fixed, so four times the batch means a quarter of

the steps per epoch. Each step still moves the weights a distance set by the

learning rate. So the run has made a quarter as many moves of the same size: it

has travelled a quarter of the distance through weight space on the same amount

of data. The gradients were better. There were far fewer of them.

The correction is to raise the learning rate so that the total distance

travelled is preserved.

The lesson stops here

4 more paragraphs to go

You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.

The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.

See the planThe contents

This is the reading half

Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.

The contents

The rest of this course

  1. 01Four Lines, and Genuinely All of It
  2. 02The Four Hops Between Storage and the Stepopening only
  3. 03The Starting Scale, and Why It Decides Everythingopening only
  4. 04One Number, and Why It Cannot Be One Numberopening only
  5. 05What a Bigger Batch Actually Buysyou are here
  6. 06The Four Numbers Worth Watchingopening only
  7. 07Three Ways a Run Breaks, and What Each One Looks Likeopening only
  8. 08Nobody Trains to Convergenceopening only

Read alongside