ContentsThe library

Many Specialists, One Model

Knowing More Without Doing More

In an ordinary model every parameter runs on every input, so what it knows and what it costs are the same number. Replacing one block with many and visiting a few separates them.

Here is an assumption so standard that it is rarely stated. When a model

processes an input, all of it runs. Every weight is read, every multiplication

performed, for every token. That is what makes the arithmetic easy to count and

it is also what ties two things together that have no business being tied.

The two quantities

How much a model knows is roughly a function of how many parameters it has. How

long it takes to produce a token is a function of how many parameters it runs.

In an ordinary model those are the same number, so there is no way to buy more

of the first without paying for the second.

FIG 1One block, or several with a chooser
The change is structural rather than mathematical. The arithmetic inside a block is unchanged; what changed is that there are now several blocks and a rule for picking one.

It is worth pausing on how unusual that coupling is. Almost nothing else works

this way. A reference book does not take longer to consult because it has more

entries; a filing system does not slow down because more files are in it. Those

things have an index, and the index is what makes the size of the collection

irrelevant to the cost of one lookup. A dense model has no index. It consults

everything, every time, and the only reason this seems normal is that it has

always been how these models were built.

Which block gets copied

Not every part of a layer is a good candidate. A modern layer has a part that

mixes positions together and a part that processes each position on its own, and

the second one holds most of the parameters.

FIG 2Where the parameters of one layer sit
Two thirds of the weights are in the block that treats each position independently. That block is both the largest target and the easiest to replicate, because positions are already handled separately inside it, so routing different positions to different copies changes nothing about how it works.

The part that mixes positions is left alone. It is smaller, and more importantly

it is the part that compares positions against each other, so sending different

positions through different copies of it would mean comparing things that had

been processed by different weights. The per-position block has no such problem:

it already handles each position in isolation, so which copy a position goes

through is nobody else's business.

Counting what a token touches

Once the copies exist, a model has two parameter counts, and reporting only one

of them is how this subject is usually misrepresented.

FIG 3What a single token actually runs
active parameters: what one token touches, which is what the time per token is proportional to
the shared parts, run by every token regardless of routing
how many copies each token is sent to, usually one or two
how many copies exist, which can be in the hundreds
the parameters held across all the copies together
The fraction in front is the whole trick. Hold a hundred copies and visit two and the expert parameters contribute a fiftieth of their count to the per-token cost, while contributing their full count to what the model has learned.
FIG 4Three models, counted both ways
parameters held, in billparameters a token touchratio
a dense model8.08.01.0
a sparse model with eigh47.013.03.6
a sparse model with many671.037.018.1
The first row is the ordinary case where the two counts are identical. The marked cell is what the arrangement is capable of: a model holding eighteen times more than it runs. Note that the third row has a smaller active count than the second despite holding fourteen times as much.

What it does not buy

Every one of those copies has to exist somewhere. A model holding six hundred

billion parameters needs memory for six hundred billion parameters whether a

token visits them or not, so it needs many times the hardware of a dense model

with the same speed per token. The arithmetic was genuinely saved. The memory

was not, and in fact it got much worse.

FIG 5Two costs as a model is made larger
0.00200.00400.00600.00800.0037.0202.8368.5534.3700.0parameters held, in billions
a dense model: everything held is runa sparse model: the visited share stays put
The flat line is why anyone does this. The thing it does not show is the memory, which follows the rising line in both cases, so a model sitting at the right-hand end of the flat line is cheap to run and expensive to house.

That trade is the shape of everything that follows. The arrangement converts a

problem of arithmetic into a problem of memory and of moving data between

machines, and whether that conversion is a good deal depends entirely on which

of those you have spare. An operator with idle memory and saturated arithmetic

gets a large win. An operator with the opposite balance gets a model that is

harder to deploy and no faster.

There is a further cost that the two parameter counts do not show. The copies

are spread across machines, so a token routed to a copy on another machine has

to be sent there and its result sent back, twice per layer. That communication

is not arithmetic and does not appear in any parameter count, and on a poorly

arranged system it is the largest single term in the time per token.

What to hold on to

In an ordinary model, knowing more means running more, because the same weights

do both jobs. Making several copies of the largest block and visiting a few of

them separates the two: parameters held can grow freely while parameters touched

per token stays flat. Quality follows the first number and speed follows the

second. What does not improve is memory, which now has to hold everything, and

the rest of this course is largely about the consequences of that.

Recap

  • A dense model uses all of itself on every input, which means its knowledge and its running cost are tied together and cannot be adjusted separately.
  • Replacing one block with several copies and sending each input to a small number of them gives a model many times the parameters at roughly unchanged arithmetic per input, which is the entire idea.
  • What this buys is arithmetic, not memory: every copy still has to be held somewhere even though most of them are idle for any given input, and that moved cost is where most of the difficulty lives.

This is the reading half

Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.

The contents

NextChoosing Which Parts Run →

The rest of this course

  1. 01Knowing More Without Doing Moreyou are here
  2. 02One Small Matrix Decides Everythingopening only
  3. 03Not Chemistry, Not Law, Not Medicineopening only
  4. 04The Favourite That Makes Itselfopening only
  5. 05Some Tokens Do Not Get Inopening only
  6. 06The Exponential in the Middle of Everythingopening only
  7. 07The Model You Must Hold and the Model You Runopening only
  8. 08The Comparison That Actually Means Somethingopening only

Read alongside