ContentsThe library

Keeping The Numbers In Range

Two Decisions That Look Like Details

Last timeResetting The Scale Partway

Which values the average is taken over, and whether the step sits before or after the block it serves, change what a normalised network can do and how deep it can be made.

The previous lesson left two questions open and treated them as details. They are

not details. Which values the average is taken over decides whether the method works

at all for sequences, and which side of the block the step sits on decides whether

the network can be made deep. Both choices were settled by experiment rather than

derivation, and both answers are now near universal.

Start with the grouping. The original method averages across the examples in the

batch. The alternative is to average across the features of a single example: take

one position in one sequence, look at the few thousand numbers describing it, and

normalise those against each other.

FIG 1
GroupingDepends on the batchAsserts is comparableWhere it is used
Across the batch, per channelyesthe examplesthe original method, image models
Across the features of one examplenothe features of a positionlanguage models
Across the features, in groupsnoeach small group of channelsimage models, small batches
Across one channel's spatial extentnothe positions in an imagestyle transfer
Across the features, no centringnothe features of a positionmost recent large models

The second column explains almost everything. Anything that depends on the batch

has to behave differently when serving, and the bottom row is where nearly every

large language model now sits. The middle row is the compromise reached in vision,

where channels are genuinely not comparable with one another but small groups of

them are.

Why sequences forced the change

The batch form has a specific problem with sequences that goes beyond the general

awkwardness. Sequences have different lengths. If the average for a channel is taken

across the batch at each position, then the later positions are averaged over fewer

examples, because the shorter sequences have ended. The estimates get noisier along

the sequence, and at positions longer than any training sequence there are no

estimates at all.

Averaging inside the example makes all of that disappear. Each position is

normalised against its own features, so the length of the sequence, the number of

sequences in the batch, and whether the model is training or serving are all

irrelevant. This is why the recurrent and attention-based models adopted it

immediately and never went back.

The lesson stops here

3 more paragraphs to go

You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.

The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.

See the planThe contents

This is the reading half

Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.

The contents

The rest of this course

  1. 01The Numbers A Machine Cannot Hold
  2. 02The Compounding Nobody Budgets Foropening only
  3. 03The Decision Made Before Anything Runsopening only
  4. 04Putting The Numbers Back Where They Belongopening only
  5. 05Two Decisions That Look Like Detailsyou are here
  6. 06The Road That Goes Roundopening only
  7. 07The Same Trouble, Running Backwardsopening only
  8. 08What A Training Curve Is Telling Youopening only

Read alongside