The Exponential in the Middle of Everything
Last timeThe Buffer That Overflows
Sparse models diverge more readily than dense ones, and the reason sits in the chooser: an exponential of unbounded scores, computed in low precision, deciding things discretely.
Sparse models fail to train more often than dense ones. Not dramatically more,
but enough that every account of building one spends time on it, and the cause
is concentrated in a single small operation.
The operation in question
The chooser produces its scores by multiplying the token representation by a
matrix. Nothing anywhere constrains how large those numbers are. Then they are
exponentiated.
An exponential is a very unforgiving function to apply to something unbounded.
In the sixteen-bit number formats that training uses, the largest representable
value is reached by an exponent somewhere in the tens, and well before that the
exponential of the largest score so dominates the sum that every other score
rounds away to nothing. At that point the chooser has silently become a function
that always picks the same copy, and the gradient flowing back through it is
either zero or infinite.
- the raw score given to copy i for token b, straight out of the chooser matrix
- a smooth stand-in for the largest score, which grows when any score grows
- the number of tokens in the batch, so the term is an average rather than a total
- a small coefficient, typically about a thousandth
- the penalty, added to the training loss alongside the balancing term
It is worth noticing why nothing else in the model pulls the scores back down.
The training objective cares only about which score is largest, because that is
what the selection reads, and the ordering of a set of numbers is unchanged by
adding a constant to all of them or by scaling them all up. So there is an
entire direction in which the chooser parameters can drift without the loss
registering any change at all, and a direction the loss is indifferent to is a
direction that will eventually be wandered along. The penalty exists to make the
loss care about that direction, and it is the only thing in the model that does.
The lesson stops here
6 more paragraphs to go
You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.
The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.
See the planThe contentsThis is the reading half
Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.
The contents