The Starting Scale, and Why It Decides Everything
Last timeGetting the Data to the Chip
Weights start as random numbers, and the one thing that matters about them is their scale: get it wrong and a deep network cannot be trained at all.
Why not zero
Weights have to start somewhere. Zero is the obvious choice and it is the one
choice that cannot work, for a reason worth understanding properly because it
recurs.
Consider one layer with several units, all of whose incoming weights are zero.
Every unit computes the same thing, which is nothing. Now run a step. The
gradient for a weight depends on the input it is attached to and on what came
back from above. Both of those are identical for every unit, because the units
are identical. So every unit receives exactly the same gradient and takes
exactly the same step, and after the step the units are still identical to one
another.
This never breaks. A layer of a thousand units initialised identically is, and
remains forever, a layer of one unit copied a thousand times. The capacity you
paid for does not exist. Randomness in the starting values is not a numerical
nicety: it is the only thing that makes the units of a layer different from one
another, and the technical name for the problem is symmetry breaking.
Biases are the exception, and can safely start at zero, because the weights
already break the symmetry.
What a layer does to the scale
So the weights are random. The question is how large. To answer it, follow what
a layer does to the size of what passes through it.
A unit's output is a sum of n terms, each one an input times a weight. If those
terms are independent with mean zero, the variance of a sum is the sum of the
variances. That single fact is the whole argument.
- how spread out a unit's output is
- the fan-in: how many inputs each unit sums over
- how spread out the starting weights are, which is the thing you choose
- how spread out the layer's input is
Read it as a statement about amplification. A layer with a fan-in of a thousand
and weights of variance one will produce outputs a thousand times more spread
out than its inputs. Stack a few of those and the numbers leave the range the
hardware can represent.
The fix falls straight out. Choose the weight variance so that the product is
one.
The lesson stops here
4 more paragraphs to go
You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.
The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.
See the planThe contentsThis is the reading half
Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.
The contents