Serving a Model to Many People
The Two Jobs That Get in Each Other's Way
Last timeServing Many Requests Together
Reading a prompt and writing an answer are different kinds of work on the same hardware. Mixed together they interfere, and the interference is felt by whoever was already mid-sentence.
A request has two quite distinct lives. First its prompt is read, all of it, in
one go. Then its answer is produced, one word at a time, with the whole model
read from memory for each word. These two activities look similar from the
outside and are almost opposites underneath.
| ms reading the prompt | ms writing the answer | arithmetic units kept bu | memory bandwidth used, p | |
|---|---|---|---|---|
| short prompt, short answ | 12 | 480 | 18 | 92 |
| short prompt, long answe | 12 | 4800 | 14 | 95 |
| long document, short ans | 1400 | 480 | 88 | 30 |
| long document, long answ | 1400 | 4800 | 25 | 88 |
Why they are not the same job
Reading a prompt processes every word in it simultaneously. The model is read
once and a great deal of arithmetic is done with it, so the processor's
arithmetic units are the thing that runs out. Producing a word processes a single
position. The model is read once again, for a tiny fraction of the arithmetic, so
the memory is the thing that runs out and most of the arithmetic capability sits
idle.
A machine saturated on arithmetic has memory bandwidth to spare, and a machine
saturated on memory has arithmetic to spare. That sounds like an opportunity, and
it is, but only if the two are arranged carefully. Done carelessly, they simply
take turns and each waits for the other.
- the time from arrival to the first word appearing
- time spent queued, before anything was done for this request
- how many words are in the prompt
- the arithmetic each prompt word requires
- how fast the arithmetic can be done
What one long prompt does to everybody
Here is the part that produces confusing incident reports. A step is a shared
event. Whatever the step is doing, every request in the group waits for the step
to finish.
So when a prompt of two thousand words is admitted and read in a single step,
that step takes perhaps a second. Every other request in the group, each of them
quietly producing a word every forty milliseconds, produces nothing at all for
that second. Their users see the text stop mid-sentence and then resume.
The lesson stops here
3 more paragraphs to go
You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.
The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.
See the planThe contentsThis is the reading half
Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.
The contents