The Decision Made Before Anything Runs
Last timeFifty Multiplications Later
The starting weights are drawn from a distribution whose width is computed from the size of the layer, because that is the only width at which a deep stack neither grows nor collapses.
The previous lesson left an equation to satisfy. The gain of a layer is its number
of inputs multiplied by the spread of its weights and by whatever the activation
contributes, and that product has to be one. Two of those three quantities are
fixed by the architecture. The third is the only free parameter, and choosing it is
what initialisation means.
Before the formula, a prior question. Why not start every weight at zero? The
network would be perfectly well defined, the first forward pass would produce
zeros, and the gradients would tell it where to go.
| step | step | state of two units in the same layer | consequence |
|---|---|---|---|
| 1 | start | identical weights | the two units are indistinguishable |
| 2 | forward | identical outputs | they contribute the same thing to every |
| 3 | backward | identical gradients | nothing in the signal distinguishes one |
| 4 | update | identical new weights | the symmetry survives the step intact |
| 5 | after a thousand steps | still identical | a layer of a thousand units has the expr |
The rule, and where it comes from
A unit in a layer adds up some number of terms, each one an input multiplied by a
weight. If the terms are independent and roughly centred, the spread of a sum of
them is the number of terms multiplied by the spread of one term. That is the only
fact needed.
So if the layer has a thousand inputs and you want the output spread to match the
input spread, each term must contribute a thousandth as much as the total should
have. The weights therefore need a spread of one over a thousand, which is a
typical size of about three hundredths. Give the same layer weights of typical size
one and its output is thirty times too large before anything else has happened.
- the spread the starting weights are drawn with
- how many inputs each unit of the layer sums over
- the correction for a rectifier discarding half of its input
The two directions want different things
There is a complication. The same weights are used in both directions. On the
forward pass each unit sums over its inputs, so the relevant count is how many
inputs the layer has. On the backward pass the gradient arriving at a unit is a sum
over everything that unit fed into, so the relevant count is how many outputs the
layer has.
When the layer is square these are the same and there is no conflict. When the
layer changes width, they are not, and no single spread satisfies both. The usual
answer is to split the difference by using the average of the two counts, which
preserves neither quantity exactly and keeps both within a modest factor. In
practice networks are mostly built from layers of similar widths, so the
disagreement is small and the compromise costs almost nothing.
The lesson stops here
3 more paragraphs to go
You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.
The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.
See the planThe contentsThis is the reading half
Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.
The contents