Putting The Numbers Back Where They Belong
Last timeWhere The Weights Start
A normalisation step subtracts the average and divides by the spread, which forces the signal back to a known size no matter what the layer before it did, and then hands two learned parameters back.
Initialisation gets the scale right once. The problem is that it cannot hold,
because training moves the weights and moving the weights moves the gain. The
obvious response, once you state the problem that way, is to stop relying on the
scale being right and simply impose it.
That is what a normalisation step does. Partway through the network, take whatever
arrived, measure its average and its spread, and transform it so that the average
is zero and the spread is one. Then carry on. Whatever the previous layer did to
the magnitude is erased.
- one value arriving at the step
- the average of the collection this value belongs to
- the spread of that same collection
- a small constant, so that a collection with no variation does not cause a division by zero
Why this stops the drift
The important property is a cancellation. Suppose training has doubled every weight
in a layer. Its outputs are then twice as large as before. But the average of those
outputs has also doubled, and so has their spread, so after subtracting the one and
dividing by the other the result is exactly what it would have been before the
weights grew.
The consequence is that the size of a layer's output no longer depends on the
magnitude of its weights at all. The weights still determine what the layer
computes, in the sense of which directions it responds to, but they no longer
determine how large the response is. Since the compounding problem is entirely about
magnitudes, removing the magnitude from the weights removes the compounding.
| step | value | arriving | arriving after the weights doubled | after normalising, both cases |
|---|---|---|---|---|
| 1 | first | 2.0 | 4.0 | -1.41 |
| 2 | second | 3.0 | 6.0 | -0.71 |
| 3 | third | 4.0 | 8.0 | 0.00 |
| 4 | fourth | 5.0 | 10.0 | 0.71 |
| 5 | fifth | 6.0 | 12.0 | 1.41 |
Giving the freedom back on purpose
Forcing every intermediate signal to have zero average and unit spread is a real
restriction. Some layers genuinely want a large output. Some want an offset, because
the activation that follows is not symmetric about zero. Most importantly, a layer
that should do nothing at all can no longer do nothing, because normalisation will
rescale its output regardless.
So the step is followed by two learned parameters per channel, a multiplier and an
offset. The network can use them to restore whatever scale and centre it actually
needs, including exactly undoing the normalisation if that is best.
The lesson stops here
3 more paragraphs to go
You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.
The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.
See the planThe contentsThis is the reading half
Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.
The contents