Three Budgets, and Only One of Them Is the Weights
A served model occupies memory in three quite separate ways, and the famous techniques each attack only one of them. Here is the accounting, done in bytes.
A model does not fit, or it fits and costs too much to run. The usual next step
is to reach for a technique with an impressive number attached to it, apply it,
and find that the thing that was too big is still too big. This happens because
there are three separate memory budgets in a served system and the famous
techniques each attack one of them.
So this lesson is accounting. No technique until the next one. The only
question here is where the bytes are, because that is what decides which of the
following seven lessons applies to your problem.
Three places the memory goes
The weights are the model itself: every number learned during training, sitting
in memory from the moment the server starts until it stops. One copy per
machine, regardless of how many people are talking to it. This is the number
everybody quotes, and it is a fixed cost.
The per-request state is different in kind. As a request is processed, every
token already read leaves something behind that has to be kept for the rest of
that request, so it does not have to be recomputed for every following token.
That state belongs to one conversation. It grows as the conversation grows, and
it is paid again for every request being served at the same time.
The working space is the arithmetic in flight: the intermediate results of the
current step, alive for a fraction of a second and then gone. It scales with
how many requests are being processed together and with how wide the model is,
and it is usually the smallest of the three, but it is not nothing and it is
the one that causes the failure that looks random.
The ratio in that picture is not fixed. Serve the same model to four users with
short prompts and the weights are almost everything. Serve it to sixty users
with long documents pasted in and the per-request state passes the weights and
keeps going. The same model, the same machine, and a different answer about
what to do.
The weights are one multiplication
Weight memory has no subtlety in it at all.
- memory the weights occupy, in bytes
- number of parameters in the model
- bytes stored per parameter
Seven billion parameters at two bytes each is fourteen gigabytes. The same model
at four bits is three and a half. There is no cleverness in the arithmetic: the
entire question is what rounding each parameter to a cruder number does to the
answers, which is the subject of the next two lessons.
It is worth noticing what this formula does not contain. It does not depend on
how many users you have, how long their prompts are, or how fast you want to
answer. That independence is exactly why the weights are the easy budget, and
why a system whose real problem is elsewhere gets no relief from attacking them.
The state you keep per request
The per-request state has more factors, and that is why it grows faster than
people expect.
- total per-request state, in bytes
- number of layers in the model
- size of the state each layer keeps per token
- tokens in the conversation so far
- requests being served at the same time
- bytes per stored number
Put numbers in. A model with forty layers keeping eight kilobytes of state per
token per layer holds about a third of a megabyte for every token of context.
A four thousand token conversation is therefore above a gigabyte, for one user.
Thirty users at once is more than thirty gigabytes, which is larger than many
of the models people are trying to shrink.
Three things follow from the shape of those lines. Long context is not a
feature you add for free: it is a memory cost paid per user, every second they
are connected. Concurrency and context trade directly against each other, so
the same machine serves many short conversations or a few long ones. And
because this budget is linear while the weights are constant, every system
crosses over at some traffic level, after which shrinking the model stops being
the thing that helps.
Which budget each technique attacks
With the accounting done, the technique list stops being a bag of tricks and
becomes a lookup.
| Weights cut, per cent | Per-request state cut, p | Working space cut, per c | Typical quality cost, po | |
|---|---|---|---|---|
| Sixteen bits to eight, w | 50 | 0 | 0 | 0 |
| Sixteen bits to four, we | 75 | 0 | 0 | 2 |
| Scattered weight removal | 0 | 0 | 0 | 1 |
| Whole structures removed | 30 | 30 | 30 | 4 |
| Smaller model trained on | 75 | 75 | 75 | 5 |
| Per-request state in eig | 0 | 50 | 0 | 1 |
Read the table as a matching problem rather than a ranking. If the weights are
the budget that does not fit, the top two rows apply and they are cheap. If the
per-request state is the budget that does not fit, the only rows that help are
the last three, and two of them require retraining. If what you want is speed
rather than room, note that only rows that remove whole structures or change
the model itself move the arithmetic, since fewer bits mainly moves bytes.
That is the discipline this whole course rests on. Measure the three budgets on
your own serving machine under your own traffic first. The number that is large
tells you which of the next seven lessons you are actually reading, and the
technique everybody is talking about is frequently an answer to somebody else's
budget.
Recap
- The weights are a fixed cost you pay once per machine. The per-request state is a cost you pay again for every user, and it is the one that usually runs you out of room.
- Weight memory is parameters times bytes per parameter, and nothing else. Per-request memory grows with context length and with how many requests you serve at once.
- Before choosing a technique, work out which budget is large. Halving the weights of a system that ran out of room on cache buys you almost nothing.
This is the reading half
Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.
The contents