ContentsThe library

Many Specialists, One Model

One Small Matrix Decides Everything

Last timeCapacity You Do Not Have to Run

A single matrix scores every copy for every token, the highest few are visited, and their outputs are blended by those same scores. That last detail is the only reason the chooser learns anything.

The copies exist. Something has to decide where each token goes, and that

something turns out to be the smallest part of the whole arrangement and the

source of nearly all its difficulties.

The decision itself

A single matrix. The token representation goes in, one number per copy comes

out, those numbers are turned into a distribution, and the highest few win. That

is the entire mechanism, and on the arithmetic it is a rounding error next to

the copies themselves.

FIG 1One token through a layer with copies
Only the second and the last boxes involve the chooser, and between them sits all the work. The chooser costs nothing to run and determines everything about what the layer does and where the time goes.
FIG 2What the layer outputs
the set of copies chosen for this token, usually one or two out of many
the score the chooser gave copy i, renormalised across the chosen set so the weights sum to one
what copy i produces for this token, which is ordinary dense arithmetic inside
the output of the layer, the same shape a single dense block would have produced
Note that the score appears twice in the story: once in deciding which copies are in the chosen set, and once as a multiplier on what they produce. The first use is discrete and the second is not, and that difference is why any of this can be trained.

Why the weights are not decoration

Choosing the highest score is a step function. Nudge the scores slightly and

nothing changes until one of them crosses another, at which point the output

jumps. A function like that gives the chooser no information about how to

improve, because almost everywhere its derivative is zero.

Multiplying the chosen copy output by its own score fixes this. The score is now

part of the output, continuously, so if raising that copy contribution would

have improved the result, the training signal says so directly and the score

rises. The chooser learns from the copies it already picked. It learns nothing

about the ones it did not pick, which is a limitation worth remembering, and it

explains a great deal about the failure in two lessons' time.

It is worth being precise about how small the chooser is, because the

asymmetry is what makes the whole arrangement possible. A single copy in a large

model might hold tens of millions of numbers. The chooser for a layer with sixty

four copies holds the width of the representation times sixty four, which is

perhaps a hundred thousand. So a part of the model that is one percent of one

percent of the size of what it governs makes the decision that determines where

all the time goes. Nothing else in the model has that ratio of influence to

size, and it explains why so much attention is paid to a component that could

be described in a sentence.

The lesson stops here

4 more paragraphs to go

You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.

The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.

See the planThe contents

This is the reading half

Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.

The contents

The rest of this course

  1. 01Knowing More Without Doing More
  2. 02One Small Matrix Decides Everythingyou are here
  3. 03Not Chemistry, Not Law, Not Medicineopening only
  4. 04The Favourite That Makes Itselfopening only
  5. 05Some Tokens Do Not Get Inopening only
  6. 06The Exponential in the Middle of Everythingopening only
  7. 07The Model You Must Hold and the Model You Runopening only
  8. 08The Comparison That Actually Means Somethingopening only

Read alongside