ContentsThe library

Making a Model Smaller

Three Budgets, and Only One of Them Is the Weights

A served model occupies memory in three quite separate ways, and the famous techniques each attack only one of them. Here is the accounting, done in bytes.

A model does not fit, or it fits and costs too much to run. The usual next step

is to reach for a technique with an impressive number attached to it, apply it,

and find that the thing that was too big is still too big. This happens because

there are three separate memory budgets in a served system and the famous

techniques each attack one of them.

So this lesson is accounting. No technique until the next one. The only

question here is where the bytes are, because that is what decides which of the

following seven lessons applies to your problem.

Three places the memory goes

The weights are the model itself: every number learned during training, sitting

in memory from the moment the server starts until it stops. One copy per

machine, regardless of how many people are talking to it. This is the number

everybody quotes, and it is a fixed cost.

The per-request state is different in kind. As a request is processed, every

token already read leaves something behind that has to be kept for the rest of

that request, so it does not have to be recomputed for every following token.

That state belongs to one conversation. It grows as the conversation grows, and

it is paid again for every request being served at the same time.

The working space is the arithmetic in flight: the intermediate results of the

current step, alive for a fraction of a second and then gone. It scales with

how many requests are being processed together and with how wide the model is,

and it is usually the smallest of the three, but it is not nothing and it is

the one that causes the failure that looks random.

FIG 1Eighty gigabytes on one serving machine
A twenty-six billion parameter model in sixteen-bit format, serving twenty-four concurrent requests at a few thousand tokens each. Halving the weights here frees twenty-six gigabytes. Halving the per-request state frees ten. Which of those is the right move depends entirely on whether you are adding users or adding context.

The ratio in that picture is not fixed. Serve the same model to four users with

short prompts and the weights are almost everything. Serve it to sixty users

with long documents pasted in and the per-request state passes the weights and

keeps going. The same model, the same machine, and a different answer about

what to do.

The weights are one multiplication

Weight memory has no subtlety in it at all.

FIG 2Bytes held by the weights
memory the weights occupy, in bytes
number of parameters in the model
bytes stored per parameter
Two factors, and only one of them is yours to change once the model is chosen. Sixteen-bit numbers mean two bytes each; eight-bit means one; four-bit means a half. Everything in the next three lessons is an attack on b, and the saving is exactly proportional because the formula has no other terms.

Seven billion parameters at two bytes each is fourteen gigabytes. The same model

at four bits is three and a half. There is no cleverness in the arithmetic: the

entire question is what rounding each parameter to a cruder number does to the

answers, which is the subject of the next two lessons.

It is worth noticing what this formula does not contain. It does not depend on

how many users you have, how long their prompts are, or how fast you want to

answer. That independence is exactly why the weights are the easy budget, and

why a system whose real problem is elsewhere gets no relief from attacking them.

The state you keep per request

The per-request state has more factors, and that is why it grows faster than

people expect.

FIG 3Bytes held by one batch of requests
total per-request state, in bytes
number of layers in the model
size of the state each layer keeps per token
tokens in the conversation so far
requests being served at the same time
bytes per stored number
The two is because each layer keeps two of these per token. The thing to see is that six factors multiply, and two of them, the context length and the number of concurrent requests, are chosen by your users and your traffic rather than by you.

Put numbers in. A model with forty layers keeping eight kilobytes of state per

token per layer holds about a third of a megabyte for every token of context.

A four thousand token conversation is therefore above a gigabyte, for one user.

Thirty users at once is more than thirty gigabytes, which is larger than many

of the models people are trying to shrink.

FIG 4Per-request state against context length
0.00100.00200.00300.00400.000.08.016.024.032.0context length, thousands of tokens
1 request8 requests32 requests
Three straight lines, because this budget is linear in both context and concurrency, and the product of two things you do not control. At thirty-two thousand tokens and thirty-two concurrent requests the state is well past three hundred gigabytes, which is the real reason long-context serving is expensive.

Three things follow from the shape of those lines. Long context is not a

feature you add for free: it is a memory cost paid per user, every second they

are connected. Concurrency and context trade directly against each other, so

the same machine serves many short conversations or a few long ones. And

because this budget is linear while the weights are constant, every system

crosses over at some traffic level, after which shrinking the model stops being

the thing that helps.

Which budget each technique attacks

With the accounting done, the technique list stops being a bag of tricks and

becomes a lookup.

FIG 5What each technique actually reduces
Weights cut, per centPer-request state cut, pWorking space cut, per cTypical quality cost, po
Sixteen bits to eight, w50000
Sixteen bits to four, we75002
Scattered weight removal0001
Whole structures removed3030304
Smaller model trained on7575755
Per-request state in eig05001
The marked row is the one worth staring at. Removing half the individual weights of a layer and writing zeros in their place saves nothing at all, because the zeros still occupy their bytes and the arithmetic still multiplies by them. It is reported as a compression technique and on ordinary hardware it compresses nothing, which is the subject of a later lesson.

Read the table as a matching problem rather than a ranking. If the weights are

the budget that does not fit, the top two rows apply and they are cheap. If the

per-request state is the budget that does not fit, the only rows that help are

the last three, and two of them require retraining. If what you want is speed

rather than room, note that only rows that remove whole structures or change

the model itself move the arithmetic, since fewer bits mainly moves bytes.

That is the discipline this whole course rests on. Measure the three budgets on

your own serving machine under your own traffic first. The number that is large

tells you which of the next seven lessons you are actually reading, and the

technique everybody is talking about is frequently an answer to somebody else's

budget.

Recap

  • The weights are a fixed cost you pay once per machine. The per-request state is a cost you pay again for every user, and it is the one that usually runs you out of room.
  • Weight memory is parameters times bytes per parameter, and nothing else. Per-request memory grows with context length and with how many requests you serve at once.
  • Before choosing a technique, work out which budget is large. Halving the weights of a system that ran out of room on cache buys you almost nothing.

This is the reading half

Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.

The contents

NextKeeping Fewer Bits →

The rest of this course

  1. 01Three Budgets, and Only One of Them Is the Weightsyou are here
  2. 02Round It, Store the Integer, Multiply It Backopening only
  3. 03Most of the Model Does Not Care, and You Have to Find the Part That Doesopening only
  4. 04Round the Finished Model, or Train One That Expects Itopening only
  5. 05A Zero Still Occupies Its Place in the Rowopening only
  6. 06The Wrong Answers Are Where the Teaching Isopening only
  7. 07The Average Moved Half a Point and the Model Cannot Add Any Moreopening only
  8. 08The Savings Multiply and So Does the Damageopening only

Read alongside