ContentsThe library

Keeping The Numbers In Range

The Compounding Nobody Budgets For

Last timeWhat A Machine Can Hold

Each layer multiplies the size of the signal by some factor, and a factor repeated fifty times is either enormous or nothing, so only a gain of almost exactly one survives depth.

Take a network and ignore everything it is for. Do not ask what it classifies or

predicts. Ask only one question about each layer: how large is what comes out

compared with what went in? That ratio is the layer's gain, and it is the entire

subject of this lesson.

A single gain is a dull number. Its interest comes from the fact that layers are

stacked, and stacking multiplies gains. The depth that makes deep networks powerful

is the same depth that turns a small, reasonable-looking gain into a number that no

format can hold.

FIG 1Where a layer's gain comes from
the spread of the outputs of layer number ell
how many inputs each unit of the layer sums over
the spread of the weights in the layer
the spread of what arrived from the previous layer
Read this as a recipe for the gain. The output spread is the input spread multiplied by the number of inputs and by the weight spread. The gain is therefore the product of the two middle factors, and for it to be one they must satisfy a single equation together. Neither can be chosen without reference to the other, which is the whole content of the next lesson.

What compounding does to a reasonable number

Suppose the gain is 1.1. A tenth larger per layer. In isolation that is nothing: a

measurement error, a rounding, something you would not bother to report. Across

fifty layers it is a factor of a hundred and seventeen. Across a hundred layers it

is nearly fourteen thousand.

Now suppose the gain is 0.9. The same logic, mirrored. Fifty layers reduce the

signal to about two percent of its original size and a hundred layers to about

three parts in a hundred thousand. In a narrow number format the second of those is

already indistinguishable from zero.

FIG 2The gain of one layer, after fifty of them
0.0037.5075.00112.50150.000.91.01.01.11.1gain of a single layer
the fiftieth power of the per-layer gain
The horizontal axis spans a range most people would call the same number twice. The vertical axis spans four orders of magnitude. Everything to the left of the centre collapses towards nothing and everything to the right climbs steeply, and the only place the curve passes through one is the single point in the middle. Make the network deeper and the curve becomes steeper still, which is the sense in which depth is numerically expensive.

That curve is the reason this course exists. It also explains a historical puzzle.

Networks of two or three layers trained perfectly well for decades with weights

chosen by rough convention, because a gain of 1.2 across three layers is a factor

of 1.7 and nothing notices. The same convention at thirty layers produces a factor

of two hundred, and nothing works at all. Depth did not need a new kind of

mathematics. It needed the scaling to be taken seriously for the first time.

The lesson stops here

3 more paragraphs to go

You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.

The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.

See the planThe contents

This is the reading half

Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.

The contents

The rest of this course

  1. 01The Numbers A Machine Cannot Hold
  2. 02The Compounding Nobody Budgets Foryou are here
  3. 03The Decision Made Before Anything Runsopening only
  4. 04Putting The Numbers Back Where They Belongopening only
  5. 05Two Decisions That Look Like Detailsopening only
  6. 06The Road That Goes Roundopening only
  7. 07The Same Trouble, Running Backwardsopening only
  8. 08What A Training Curve Is Telling Youopening only

Read alongside