ContentsThe library

Many Specialists, One Model

The Exponential in the Middle of Everything

Last timeThe Buffer That Overflows

Sparse models diverge more readily than dense ones, and the reason sits in the chooser: an exponential of unbounded scores, computed in low precision, deciding things discretely.

Sparse models fail to train more often than dense ones. Not dramatically more,

but enough that every account of building one spends time on it, and the cause

is concentrated in a single small operation.

The operation in question

The chooser produces its scores by multiplying the token representation by a

matrix. Nothing anywhere constrains how large those numbers are. Then they are

exponentiated.

An exponential is a very unforgiving function to apply to something unbounded.

In the sixteen-bit number formats that training uses, the largest representable

value is reached by an exponent somewhere in the tens, and well before that the

exponential of the largest score so dominates the sum that every other score

rounds away to nothing. At that point the chooser has silently become a function

that always picks the same copy, and the gradient flowing back through it is

either zero or infinite.

FIG 1Largest score magnitude over training
0.0010.0020.0030.0040.001.0500.81000.51500.32000.0training steps
nothing constraining the scoreswith a penalty on score magnitude
The unconstrained line has no reason to stop. Nothing in the objective cares how large a score is, only which is largest, so a slow drift upward is free and continues until the arithmetic breaks. The penalised line settles into a range where the exponential is well behaved and stays there.
FIG 2The penalty on the size of the scores
the raw score given to copy i for token b, straight out of the chooser matrix
a smooth stand-in for the largest score, which grows when any score grows
the number of tokens in the batch, so the term is an average rather than a total
a small coefficient, typically about a thousandth
the penalty, added to the training loss alongside the balancing term
Squaring makes the penalty indifferent to sign and sharply increasing, so a large score is punished much harder than a moderate one. What it notably does not do is express any preference about which score is largest, so unlike the balancing term it is not in competition with the routing the model wants.

It is worth noticing why nothing else in the model pulls the scores back down.

The training objective cares only about which score is largest, because that is

what the selection reads, and the ordering of a set of numbers is unchanged by

adding a constant to all of them or by scaling them all up. So there is an

entire direction in which the chooser parameters can drift without the loss

registering any change at all, and a direction the loss is indifferent to is a

direction that will eventually be wandered along. The penalty exists to make the

loss care about that direction, and it is the only thing in the model that does.

The lesson stops here

6 more paragraphs to go

You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.

The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.

See the planThe contents

This is the reading half

Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.

The contents

The rest of this course

  1. 01Knowing More Without Doing More
  2. 02One Small Matrix Decides Everythingopening only
  3. 03Not Chemistry, Not Law, Not Medicineopening only
  4. 04The Favourite That Makes Itselfopening only
  5. 05Some Tokens Do Not Get Inopening only
  6. 06The Exponential in the Middle of Everythingyou are here
  7. 07The Model You Must Hold and the Model You Runopening only
  8. 08The Comparison That Actually Means Somethingopening only

Read alongside