The library

The mathematics underneath

Be able to explain why deep networks fall apart numerically, what initialisation, normalisation and skip connections each do about it, and how to read a training run that is going wrong

Keeping The Numbers In Range

A deep network is a long chain of multiplications, and long chains of multiplications either blow up or collapse. Almost every structural choice in a modern architecture exists to stop that happening.

8 lessons, written and corrected before you arrived. Reading them here needs no account. The first reads the whole way through; the others open and then stop, because a page nobody owns cannot tell who is reading it. Starting the course gives you your own copy, where every idea has problems standing under it and you can ask about any sentence.

Start reading

  1. 01The Numbers A Machine Cannot HoldA number format buys you a range of magnitudes and a number of significant digits, and both budgets are small enough that ordinary arithmetic walks off either end of them.
  2. 02The Compounding Nobody Budgets Foropening onlyEach layer multiplies the size of the signal by some factor, and a factor repeated fifty times is either enormous or nothing, so only a gain of almost exactly one survives depth.
  3. 03The Decision Made Before Anything Runsopening onlyThe starting weights are drawn from a distribution whose width is computed from the size of the layer, because that is the only width at which a deep stack neither grows nor collapses.
  4. 04Putting The Numbers Back Where They Belongopening onlyA normalisation step subtracts the average and divides by the spread, which forces the signal back to a known size no matter what the layer before it did, and then hands two learned parameters back.
  5. 05Two Decisions That Look Like Detailsopening onlyWhich values the average is taken over, and whether the step sits before or after the block it serves, change what a normalised network can do and how deep it can be made.
  6. 06The Road That Goes Roundopening onlyAdding a block's input back to its output makes doing nothing the default, gives the gradient a route that is never rescaled, and turns a deep stack into a collection of mostly short paths.
  7. 07The Same Trouble, Running Backwardsopening onlyThe gradient travels back through the same layers as a product of their effects, so it vanishes or explodes for identical reasons, and the remedies are different because the symptoms are.
  8. 08What A Training Curve Is Telling Youopening onlyEverything in this course arrives in practice as one line on a chart going the wrong way, so the last skill is reading that line and knowing which of the remedies to reach for.