One Small Matrix Decides Everything
Last timeCapacity You Do Not Have to Run
A single matrix scores every copy for every token, the highest few are visited, and their outputs are blended by those same scores. That last detail is the only reason the chooser learns anything.
The copies exist. Something has to decide where each token goes, and that
something turns out to be the smallest part of the whole arrangement and the
source of nearly all its difficulties.
The decision itself
A single matrix. The token representation goes in, one number per copy comes
out, those numbers are turned into a distribution, and the highest few win. That
is the entire mechanism, and on the arithmetic it is a rounding error next to
the copies themselves.
- the set of copies chosen for this token, usually one or two out of many
- the score the chooser gave copy i, renormalised across the chosen set so the weights sum to one
- what copy i produces for this token, which is ordinary dense arithmetic inside
- the output of the layer, the same shape a single dense block would have produced
Why the weights are not decoration
Choosing the highest score is a step function. Nudge the scores slightly and
nothing changes until one of them crosses another, at which point the output
jumps. A function like that gives the chooser no information about how to
improve, because almost everywhere its derivative is zero.
Multiplying the chosen copy output by its own score fixes this. The score is now
part of the output, continuously, so if raising that copy contribution would
have improved the result, the training signal says so directly and the score
rises. The chooser learns from the copies it already picked. It learns nothing
about the ones it did not pick, which is a limitation worth remembering, and it
explains a great deal about the failure in two lessons' time.
It is worth being precise about how small the chooser is, because the
asymmetry is what makes the whole arrangement possible. A single copy in a large
model might hold tens of millions of numbers. The chooser for a layer with sixty
four copies holds the width of the representation times sixty four, which is
perhaps a hundred thousand. So a part of the model that is one percent of one
percent of the size of what it governs makes the decision that determines where
all the time goes. Nothing else in the model has that ratio of influence to
size, and it explains why so much attention is paid to a component that could
be described in a sentence.
The lesson stops here
4 more paragraphs to go
You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.
The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.
See the planThe contentsThis is the reading half
Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.
The contents