ContentsThe library

Training a Model From Scratch

The Starting Scale, and Why It Decides Everything

Last timeGetting the Data to the Chip

Weights start as random numbers, and the one thing that matters about them is their scale: get it wrong and a deep network cannot be trained at all.

Why not zero

Weights have to start somewhere. Zero is the obvious choice and it is the one

choice that cannot work, for a reason worth understanding properly because it

recurs.

Consider one layer with several units, all of whose incoming weights are zero.

Every unit computes the same thing, which is nothing. Now run a step. The

gradient for a weight depends on the input it is attached to and on what came

back from above. Both of those are identical for every unit, because the units

are identical. So every unit receives exactly the same gradient and takes

exactly the same step, and after the step the units are still identical to one

another.

This never breaks. A layer of a thousand units initialised identically is, and

remains forever, a layer of one unit copied a thousand times. The capacity you

paid for does not exist. Randomness in the starting values is not a numerical

nicety: it is the only thing that makes the units of a layer different from one

another, and the technical name for the problem is symmetry breaking.

Biases are the exception, and can safely start at zero, because the weights

already break the symmetry.

What a layer does to the scale

So the weights are random. The question is how large. To answer it, follow what

a layer does to the size of what passes through it.

A unit's output is a sum of n terms, each one an input times a weight. If those

terms are independent with mean zero, the variance of a sum is the sum of the

variances. That single fact is the whole argument.

FIG 1Variance through one layer
how spread out a unit's output is
the fan-in: how many inputs each unit sums over
how spread out the starting weights are, which is the thing you choose
how spread out the layer's input is
The fan-in enters as a plain multiplier, so a wide layer amplifies by default and must be compensated for.

Read it as a statement about amplification. A layer with a fan-in of a thousand

and weights of variance one will produce outputs a thousand times more spread

out than its inputs. Stack a few of those and the numbers leave the range the

hardware can represent.

The fix falls straight out. Choose the weight variance so that the product is

one.

The lesson stops here

4 more paragraphs to go

You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.

The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.

See the planThe contents

This is the reading half

Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.

The contents

The rest of this course

  1. 01Four Lines, and Genuinely All of It
  2. 02The Four Hops Between Storage and the Stepopening only
  3. 03The Starting Scale, and Why It Decides Everythingyou are here
  4. 04One Number, and Why It Cannot Be One Numberopening only
  5. 05What a Bigger Batch Actually Buysopening only
  6. 06The Four Numbers Worth Watchingopening only
  7. 07Three Ways a Run Breaks, and What Each One Looks Likeopening only
  8. 08Nobody Trains to Convergenceopening only

Read alongside