The Same Trouble, Running Backwards
Last timeA Path Around The Layer
The gradient travels back through the same layers as a product of their effects, so it vanishes or explodes for identical reasons, and the remedies are different because the symptoms are.
Everything so far has followed a signal forwards. The gradient travels the other
way, through the same weights, and it is also a product with one factor per layer.
That is the entire reason this lesson is short: the mathematics is the same, so the
compounding conclusions carry over without modification.
What does not carry over is the practical handling. The forward pass can be fixed
structurally, by choosing the starting scale, imposing the scale partway through, or
building a route that scales nothing. The backward pass has two additional problems
that structure does not address, and both have their own standard remedy.
- what reaches the bottom of the stack
- the gradient at the output, where the loss is measured
- the product over every layer between the output and this point
- the effect of one layer on what passes through it
The two directions of failure, and how differently they present
The arithmetic is symmetric and the experience of debugging is not.
An exploding gradient announces itself. The step is enormous, the parameters land
somewhere arbitrary, and the loss rises to a large value or becomes unreadable in a
single iteration. You see it in the first place you look, which is the loss curve.
A vanishing gradient hides. The later layers continue to train normally, because
their gradients have travelled through few factors. The early layers receive almost
nothing and stay near their starting values. The loss falls for a while, as the upper
part of the network does what it can with whatever features the untrained lower part
happens to produce, and then flattens at a mediocre value. Nothing about the curve
says that half the network is not participating.
Rescaling an exploding gradient
The remedy for the loud failure is blunt and works well. Before taking a step,
measure the overall length of the gradient across every parameter in the model. If
it is above a chosen threshold, multiply the whole gradient by the factor that
brings it exactly to the threshold.
Two properties make this reasonable. The direction is preserved exactly, so the step
goes where the gradient said, just not as far. And the measurement is global, so the
relative sizes of the different parts of the gradient survive, which matters because
those relative sizes are what tells the optimiser which parameters are implicated.
Clipping each parameter separately would destroy that and is not what is meant by
clipping.
The lesson stops here
3 more paragraphs to go
You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.
The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.
See the planThe contentsThis is the reading half
Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.
The contents