ContentsThe library

Keeping The Numbers In Range

Putting The Numbers Back Where They Belong

Last timeWhere The Weights Start

A normalisation step subtracts the average and divides by the spread, which forces the signal back to a known size no matter what the layer before it did, and then hands two learned parameters back.

Initialisation gets the scale right once. The problem is that it cannot hold,

because training moves the weights and moving the weights moves the gain. The

obvious response, once you state the problem that way, is to stop relying on the

scale being right and simply impose it.

That is what a normalisation step does. Partway through the network, take whatever

arrived, measure its average and its spread, and transform it so that the average

is zero and the spread is one. Then carry on. Whatever the previous layer did to

the magnitude is erased.

FIG 1The normalisation step
one value arriving at the step
the average of the collection this value belongs to
the spread of that same collection
a small constant, so that a collection with no variation does not cause a division by zero
There is nothing learned in this expression and nothing approximate about it. The output has an average of zero and a spread of one as a matter of arithmetic, not as a tendency. The only design decisions are which collection the average is taken over, which is the subject of the next lesson, and how small the constant should be.

Why this stops the drift

The important property is a cancellation. Suppose training has doubled every weight

in a layer. Its outputs are then twice as large as before. But the average of those

outputs has also doubled, and so has their spread, so after subtracting the one and

dividing by the other the result is exactly what it would have been before the

weights grew.

The consequence is that the size of a layer's output no longer depends on the

magnitude of its weights at all. The weights still determine what the layer

computes, in the sense of which directions it responds to, but they no longer

determine how large the response is. Since the compounding problem is entirely about

magnitudes, removing the magnitude from the weights removes the compounding.

FIG 2Five values through the step, before and after the weights grew
stepvaluearrivingarriving after the weights doubledafter normalising, both cases
1first2.04.0-1.41
2second3.06.0-0.71
3third4.08.00.00
4fourth5.010.00.71
5fifth6.012.01.41
5 steps
The two input columns differ by a factor of two and the output column is identical for both. That identity is the mechanism. Note also what is preserved: the relative ordering and the relative spacing of the five values survive untouched, so no information about which example is which has been discarded. Only the scale has gone.

Giving the freedom back on purpose

Forcing every intermediate signal to have zero average and unit spread is a real

restriction. Some layers genuinely want a large output. Some want an offset, because

the activation that follows is not symmetric about zero. Most importantly, a layer

that should do nothing at all can no longer do nothing, because normalisation will

rescale its output regardless.

So the step is followed by two learned parameters per channel, a multiplier and an

offset. The network can use them to restore whatever scale and centre it actually

needs, including exactly undoing the normalisation if that is best.

The lesson stops here

3 more paragraphs to go

You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.

The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.

See the planThe contents

This is the reading half

Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.

The contents

The rest of this course

  1. 01The Numbers A Machine Cannot Hold
  2. 02The Compounding Nobody Budgets Foropening only
  3. 03The Decision Made Before Anything Runsopening only
  4. 04Putting The Numbers Back Where They Belongyou are here
  5. 05Two Decisions That Look Like Detailsopening only
  6. 06The Road That Goes Roundopening only
  7. 07The Same Trouble, Running Backwardsopening only
  8. 08What A Training Curve Is Telling Youopening only

Read alongside