Knowing More Without Doing More
In an ordinary model every parameter runs on every input, so what it knows and what it costs are the same number. Replacing one block with many and visiting a few separates them.
Here is an assumption so standard that it is rarely stated. When a model
processes an input, all of it runs. Every weight is read, every multiplication
performed, for every token. That is what makes the arithmetic easy to count and
it is also what ties two things together that have no business being tied.
The two quantities
How much a model knows is roughly a function of how many parameters it has. How
long it takes to produce a token is a function of how many parameters it runs.
In an ordinary model those are the same number, so there is no way to buy more
of the first without paying for the second.
It is worth pausing on how unusual that coupling is. Almost nothing else works
this way. A reference book does not take longer to consult because it has more
entries; a filing system does not slow down because more files are in it. Those
things have an index, and the index is what makes the size of the collection
irrelevant to the cost of one lookup. A dense model has no index. It consults
everything, every time, and the only reason this seems normal is that it has
always been how these models were built.
Which block gets copied
Not every part of a layer is a good candidate. A modern layer has a part that
mixes positions together and a part that processes each position on its own, and
the second one holds most of the parameters.
The part that mixes positions is left alone. It is smaller, and more importantly
it is the part that compares positions against each other, so sending different
positions through different copies of it would mean comparing things that had
been processed by different weights. The per-position block has no such problem:
it already handles each position in isolation, so which copy a position goes
through is nobody else's business.
Counting what a token touches
Once the copies exist, a model has two parameter counts, and reporting only one
of them is how this subject is usually misrepresented.
- active parameters: what one token touches, which is what the time per token is proportional to
- the shared parts, run by every token regardless of routing
- how many copies each token is sent to, usually one or two
- how many copies exist, which can be in the hundreds
- the parameters held across all the copies together
| parameters held, in bill | parameters a token touch | ratio | |
|---|---|---|---|
| a dense model | 8.0 | 8.0 | 1.0 |
| a sparse model with eigh | 47.0 | 13.0 | 3.6 |
| a sparse model with many | 671.0 | 37.0 | 18.1 |
What it does not buy
Every one of those copies has to exist somewhere. A model holding six hundred
billion parameters needs memory for six hundred billion parameters whether a
token visits them or not, so it needs many times the hardware of a dense model
with the same speed per token. The arithmetic was genuinely saved. The memory
was not, and in fact it got much worse.
That trade is the shape of everything that follows. The arrangement converts a
problem of arithmetic into a problem of memory and of moving data between
machines, and whether that conversion is a good deal depends entirely on which
of those you have spare. An operator with idle memory and saturated arithmetic
gets a large win. An operator with the opposite balance gets a model that is
harder to deploy and no faster.
There is a further cost that the two parameter counts do not show. The copies
are spread across machines, so a token routed to a copy on another machine has
to be sent there and its result sent back, twice per layer. That communication
is not arithmetic and does not appear in any parameter count, and on a poorly
arranged system it is the largest single term in the time per token.
What to hold on to
In an ordinary model, knowing more means running more, because the same weights
do both jobs. Making several copies of the largest block and visiting a few of
them separates the two: parameters held can grow freely while parameters touched
per token stays flat. Quality follows the first number and speed follows the
second. What does not improve is memory, which now has to hold everything, and
the rest of this course is largely about the consequences of that.
Recap
- A dense model uses all of itself on every input, which means its knowledge and its running cost are tied together and cannot be adjusted separately.
- Replacing one block with several copies and sending each input to a small number of them gives a model many times the parameters at roughly unchanged arithmetic per input, which is the entire idea.
- What this buys is arithmetic, not memory: every copy still has to be held somewhere even though most of them are idle for any given input, and that moved cost is where most of the difficulty lives.
This is the reading half
Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.
The contents