Serving a Model to Many People
Reading the Weights Once for Everybody
Last timeTraffic Does Not Arrive Evenly
The expensive part of producing a word is reading the model, and the reading serves everyone in the group at once. Group size is the single largest lever in serving, and it is a trade.
Producing the next word requires reading every number in the model out of memory
and into the processor. On a large model that is tens of gigabytes of reading per
word, and it is the slowest thing in the whole operation. The arithmetic
performed on those numbers, which is what the hardware is nominally for, finishes
long before the next batch of numbers has arrived.
That single fact is the foundation of everything in this lesson. If you have one
request in flight, you read all those numbers to produce one word. If you have
thirty-two, you read exactly the same numbers once and produce thirty-two words.
- the time charged to one request for one word of its answer
- the time to read the whole model out of memory, paid once per step
- how many requests are in the group
- the shared cost, divided among everyone present
- the per-request work, which nobody can share
Where the gains stop
The first term is enormous and the second is small, which is why the early gains
are dramatic. Going from one request to eight can increase total output nearly
eightfold. Going from sixty-four to one hundred and twenty-eight might gain
twenty percent, because by then the shared read is no longer what you are
waiting for.
What caps the group
The group cannot simply be made huge, and the limit is memory rather than
compute. Each request in flight holds its own record of the conversation so far,
which is consulted at every step and which grows by a fixed amount with each new
word. The model itself occupies most of the available memory, and whatever is
left is what the records share.
So the cap is a division: the free memory divided by the memory one request's
record needs. That denominator grows as conversations lengthen, which means the
same service can hold a group of sixty-four short requests or eight long ones,
and the group size that worked yesterday can fail today purely because the
traffic got wordier.
This has a practical consequence that catches people out. A configured maximum
group size is not a promise, it is a ceiling that the memory may refuse to honour
at any moment. A service that runs happily at a group of sixty-four for weeks can
begin rejecting arrivals the day a few users start pasting long documents in,
because those few requests consume the memory that forty others were sharing.
The lesson stops here
4 more paragraphs to go
You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.
The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.
See the planThe contentsThis is the reading half
Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.
The contents