Two Decisions That Look Like Details
Last timeResetting The Scale Partway
Which values the average is taken over, and whether the step sits before or after the block it serves, change what a normalised network can do and how deep it can be made.
The previous lesson left two questions open and treated them as details. They are
not details. Which values the average is taken over decides whether the method works
at all for sequences, and which side of the block the step sits on decides whether
the network can be made deep. Both choices were settled by experiment rather than
derivation, and both answers are now near universal.
Start with the grouping. The original method averages across the examples in the
batch. The alternative is to average across the features of a single example: take
one position in one sequence, look at the few thousand numbers describing it, and
normalise those against each other.
| Grouping | Depends on the batch | Asserts is comparable | Where it is used |
|---|---|---|---|
| Across the batch, per channel | yes | the examples | the original method, image models |
| Across the features of one example | no | the features of a position | language models |
| Across the features, in groups | no | each small group of channels | image models, small batches |
| Across one channel's spatial extent | no | the positions in an image | style transfer |
| Across the features, no centring | no | the features of a position | most recent large models |
The second column explains almost everything. Anything that depends on the batch
has to behave differently when serving, and the bottom row is where nearly every
large language model now sits. The middle row is the compromise reached in vision,
where channels are genuinely not comparable with one another but small groups of
them are.
Why sequences forced the change
The batch form has a specific problem with sequences that goes beyond the general
awkwardness. Sequences have different lengths. If the average for a channel is taken
across the batch at each position, then the later positions are averaged over fewer
examples, because the shorter sequences have ended. The estimates get noisier along
the sequence, and at positions longer than any training sequence there are no
estimates at all.
Averaging inside the example makes all of that disappear. Each position is
normalised against its own features, so the length of the sequence, the number of
sequences in the batch, and whether the model is training or serving are all
irrelevant. This is why the recurrent and attention-based models adopted it
immediately and never went back.
The lesson stops here
3 more paragraphs to go
You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.
The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.
See the planThe contentsThis is the reading half
Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.
The contents