ContentsThe library

Keeping The Numbers In Range

The Same Trouble, Running Backwards

Last timeA Path Around The Layer

The gradient travels back through the same layers as a product of their effects, so it vanishes or explodes for identical reasons, and the remedies are different because the symptoms are.

Everything so far has followed a signal forwards. The gradient travels the other

way, through the same weights, and it is also a product with one factor per layer.

That is the entire reason this lesson is short: the mathematics is the same, so the

compounding conclusions carry over without modification.

What does not carry over is the practical handling. The forward pass can be fixed

structurally, by choosing the starting scale, imposing the scale partway through, or

building a route that scales nothing. The backward pass has two additional problems

that structure does not address, and both have their own standard remedy.

FIG 1The gradient at the bottom of the stack
what reaches the bottom of the stack
the gradient at the output, where the loss is measured
the product over every layer between the output and this point
the effect of one layer on what passes through it
Compare this with the forward recursion from the second lesson and the correspondence is exact. One factor per layer, multiplied together, with the typical factor deciding whether the product grows or collapses. The quantities are different and the structure is identical, which is why a network whose forward signal is well behaved usually has a well behaved gradient too.

The two directions of failure, and how differently they present

The arithmetic is symmetric and the experience of debugging is not.

An exploding gradient announces itself. The step is enormous, the parameters land

somewhere arbitrary, and the loss rises to a large value or becomes unreadable in a

single iteration. You see it in the first place you look, which is the loss curve.

A vanishing gradient hides. The later layers continue to train normally, because

their gradients have travelled through few factors. The early layers receive almost

nothing and stay near their starting values. The loss falls for a while, as the upper

part of the network does what it can with whatever features the untrained lower part

happens to produce, and then flattens at a mediocre value. Nothing about the curve

says that half the network is not participating.

FIG 2Both failures, over eighty layers
0.00625.001250.001875.002500.000.020.040.060.080.0layers travelled backwards
per-layer factor 0.9, the vanishing caseper-layer factor 1.1, the exploding case
Two curves from factors that differ by a fifth. After eighty layers they are separated by about twelve orders of magnitude. Note which one leaves the picture first: the exploding case is obvious almost immediately, while the vanishing case approaches the bottom of the range gradually and is still producing plausible small numbers for a long while before it produces nothing.

Rescaling an exploding gradient

The remedy for the loud failure is blunt and works well. Before taking a step,

measure the overall length of the gradient across every parameter in the model. If

it is above a chosen threshold, multiply the whole gradient by the factor that

brings it exactly to the threshold.

Two properties make this reasonable. The direction is preserved exactly, so the step

goes where the gradient said, just not as far. And the measurement is global, so the

relative sizes of the different parts of the gradient survive, which matters because

those relative sizes are what tells the optimiser which parameters are implicated.

Clipping each parameter separately would destroy that and is not what is meant by

clipping.

The lesson stops here

3 more paragraphs to go

You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.

The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.

See the planThe contents

This is the reading half

Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.

The contents

The rest of this course

  1. 01The Numbers A Machine Cannot Hold
  2. 02The Compounding Nobody Budgets Foropening only
  3. 03The Decision Made Before Anything Runsopening only
  4. 04Putting The Numbers Back Where They Belongopening only
  5. 05Two Decisions That Look Like Detailsopening only
  6. 06The Road That Goes Roundopening only
  7. 07The Same Trouble, Running Backwardsyou are here
  8. 08What A Training Curve Is Telling Youopening only

Read alongside