ContentsThe library

Keeping The Numbers In Range

The Road That Goes Round

Last timeBefore Or After, And Across What

Adding a block's input back to its output makes doing nothing the default, gives the gradient a route that is never rescaled, and turns a deep stack into a collection of mostly short paths.

The result that produced this lesson was not a theoretical worry. It was a

measurement that made no sense. A network of fifty-six plain layers had higher error

on its own training data than one of twenty layers, trained the same way on the same

data.

That ordering should be impossible. The deeper network contains the shallower one:

set the extra thirty-six layers to reproduce their inputs exactly and the two

networks compute the same function. So the deeper one can match the shallower one

and cannot be worse. Unless the optimiser cannot find that setting, which was the

conclusion, and which reframes depth as a problem of reachability rather than

capacity.

FIG 1The block, with both routes
The addition at the centre is the entire mechanism and it has no parameters. Everything in this lesson follows from the fact that one of the two routes into that addition is empty, so there is a way from the input of the network to its output that no weight touches.

Why the gradient stops vanishing

The forward argument is about reachability. The backward argument is arithmetic and

it is short.

FIG 2What the addition does to the derivative
the rate of change with respect to what arrives at the block
what the block puts out: what arrived, plus what the branch computed
the contribution of the route around, which is exactly one and cannot be anything else
the contribution of the branch, which may be anything including nearly zero
The one is the whole content of the figure. In a plain stack the gradient reaching the bottom is a product of one factor per layer, and if those factors average below one the product vanishes with depth. Here each factor is one plus something, so the product cannot collapse towards zero unless the branch terms conspire to cancel the ones, which does not happen by accident.

A stack of a hundred plain layers multiplies a hundred numbers together to get the

gradient at the bottom. If those numbers average 0.9, which is entirely ordinary,

the result is about three parts in a hundred thousand. A stack of a hundred residual

blocks has, among its routes, one that multiplies nothing at all, so the gradient at

the bottom is at least the gradient at the top.

The lesson stops here

3 more paragraphs to go

You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.

The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.

See the planThe contents

This is the reading half

Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.

The contents

The rest of this course

  1. 01The Numbers A Machine Cannot Hold
  2. 02The Compounding Nobody Budgets Foropening only
  3. 03The Decision Made Before Anything Runsopening only
  4. 04Putting The Numbers Back Where They Belongopening only
  5. 05Two Decisions That Look Like Detailsopening only
  6. 06The Road That Goes Roundyou are here
  7. 07The Same Trouble, Running Backwardsopening only
  8. 08What A Training Curve Is Telling Youopening only

Read alongside