ContentsThe library

Keeping The Numbers In Range

The Decision Made Before Anything Runs

Last timeFifty Multiplications Later

The starting weights are drawn from a distribution whose width is computed from the size of the layer, because that is the only width at which a deep stack neither grows nor collapses.

The previous lesson left an equation to satisfy. The gain of a layer is its number

of inputs multiplied by the spread of its weights and by whatever the activation

contributes, and that product has to be one. Two of those three quantities are

fixed by the architecture. The third is the only free parameter, and choosing it is

what initialisation means.

Before the formula, a prior question. Why not start every weight at zero? The

network would be perfectly well defined, the first forward pass would produce

zeros, and the gradients would tell it where to go.

FIG 1What happens if every weight starts the same
stepstepstate of two units in the same layerconsequence
1startidentical weightsthe two units are indistinguishable
2forwardidentical outputsthey contribute the same thing to every
3backwardidentical gradientsnothing in the signal distinguishes one
4updateidentical new weightsthe symmetry survives the step intact
5after a thousand stepsstill identicala layer of a thousand units has the expr
5 steps
The symmetry is preserved by every operation in the procedure, so it can never break on its own. Randomness at the start is therefore structural rather than cosmetic. It is what assigns the units different jobs, and no amount of training can do that job afterwards.

The rule, and where it comes from

A unit in a layer adds up some number of terms, each one an input multiplied by a

weight. If the terms are independent and roughly centred, the spread of a sum of

them is the number of terms multiplied by the spread of one term. That is the only

fact needed.

So if the layer has a thousand inputs and you want the output spread to match the

input spread, each term must contribute a thousandth as much as the total should

have. The weights therefore need a spread of one over a thousand, which is a

typical size of about three hundredths. Give the same layer weights of typical size

one and its output is thirty times too large before anything else has happened.

FIG 2The starting scale for a rectifier network
the spread the starting weights are drawn with
how many inputs each unit of the layer sums over
the correction for a rectifier discarding half of its input
This is the single most widely used line of code in deep learning that nobody writes by hand any more. The two in the numerator is the rectifier correction. Without it the rule preserves the weighted sums and then loses a factor of about 0.71 at every activation, which at thirty layers is a factor of three thousand lost in total.

The two directions want different things

There is a complication. The same weights are used in both directions. On the

forward pass each unit sums over its inputs, so the relevant count is how many

inputs the layer has. On the backward pass the gradient arriving at a unit is a sum

over everything that unit fed into, so the relevant count is how many outputs the

layer has.

When the layer is square these are the same and there is no conflict. When the

layer changes width, they are not, and no single spread satisfies both. The usual

answer is to split the difference by using the average of the two counts, which

preserves neither quantity exactly and keeps both within a modest factor. In

practice networks are mostly built from layers of similar widths, so the

disagreement is small and the compromise costs almost nothing.

The lesson stops here

3 more paragraphs to go

You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.

The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.

See the planThe contents

This is the reading half

Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.

The contents

The rest of this course

  1. 01The Numbers A Machine Cannot Hold
  2. 02The Compounding Nobody Budgets Foropening only
  3. 03The Decision Made Before Anything Runsyou are here
  4. 04Putting The Numbers Back Where They Belongopening only
  5. 05Two Decisions That Look Like Detailsopening only
  6. 06The Road That Goes Roundopening only
  7. 07The Same Trouble, Running Backwardsopening only
  8. 08What A Training Curve Is Telling Youopening only

Read alongside