ContentsThe library

Serving a Model to Many People

Reading the Weights Once for Everybody

Last timeTraffic Does Not Arrive Evenly

The expensive part of producing a word is reading the model, and the reading serves everyone in the group at once. Group size is the single largest lever in serving, and it is a trade.

Producing the next word requires reading every number in the model out of memory

and into the processor. On a large model that is tens of gigabytes of reading per

word, and it is the slowest thing in the whole operation. The arithmetic

performed on those numbers, which is what the hardware is nominally for, finishes

long before the next batch of numbers has arrived.

That single fact is the foundation of everything in this lesson. If you have one

request in flight, you read all those numbers to produce one word. If you have

thirty-two, you read exactly the same numbers once and produce thirty-two words.

FIG 1What one request pays per word
the time charged to one request for one word of its answer
the time to read the whole model out of memory, paid once per step
how many requests are in the group
the shared cost, divided among everyone present
the per-request work, which nobody can share
Two terms, and they behave completely differently. The first falls away as the group grows and the second does not, so efficiency improves quickly at first and then stops improving once the second term dominates.

Where the gains stop

The first term is enormous and the second is small, which is why the early gains

are dramatic. Going from one request to eight can increase total output nearly

eightfold. Going from sixty-four to one hundred and twenty-eight might gain

twenty percent, because by then the shared read is no longer what you are

waiting for.

FIG 2Total output against group size, for three sizes of per-request work
0.0010.0020.0030.0040.001.016.832.548.364.0requests served together
small per-request workmoderatelarge, as with long conversations
Every curve bends, and the one that bends earliest is the one with the most unshareable work per request. This is why the right group size for short prompts is much larger than for long conversations, and why a single configured number is wrong for mixed traffic.

What caps the group

The group cannot simply be made huge, and the limit is memory rather than

compute. Each request in flight holds its own record of the conversation so far,

which is consulted at every step and which grows by a fixed amount with each new

word. The model itself occupies most of the available memory, and whatever is

left is what the records share.

So the cap is a division: the free memory divided by the memory one request's

record needs. That denominator grows as conversations lengthen, which means the

same service can hold a group of sixty-four short requests or eight long ones,

and the group size that worked yesterday can fail today purely because the

traffic got wordier.

This has a practical consequence that catches people out. A configured maximum

group size is not a promise, it is a ceiling that the memory may refuse to honour

at any moment. A service that runs happily at a group of sixty-four for weeks can

begin rejecting arrivals the day a few users start pasting long documents in,

because those few requests consume the memory that forty others were sharing.

The lesson stops here

4 more paragraphs to go

You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.

The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.

See the planThe contents

This is the reading half

Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.

The contents

The rest of this course

  1. 01Mostly Waiting
  2. 02The Average Minute Does Not Existopening only
  3. 03Reading the Weights Once for Everybodyyou are here
  4. 04The Two Jobs That Get in Each Other's Wayopening only
  5. 05Where the Complaints Actually Liveopening only
  6. 06Saying No While You Still Canopening only
  7. 07Help That Arrives Nine Minutes Lateopening only
  8. 08A Number You Can Be Held Toopening only

Read alongside