Reading and Writing Are Different Jobs
Last timeMany Requests at Once
Taking in a prompt and producing a reply obey opposite limits, want opposite settings, and interfere with each other when run on the same machine.
Every request to a model does two quite different things, and almost every
confusing measurement in serving comes from averaging over the difference.
What the two phases are
First the prompt is taken in. Every position of it is processed in a single pass
through the weights, which produces the kept summaries for all of those positions
and, at the end, the first word of the reply. Then the reply is produced, one word
per pass, each pass reading the whole store again.
The user experiences these as two separate things: how long before anything
appears, and how quickly it then arrives. They are set by different phases and
limited by different resources.
- how many requests are grouped into the pass
- how many positions of prompt are being taken in
- operations per byte while reading a prompt, since one reading of the weights serves every position at once
- operations per byte while producing words, where each request contributes only one position
The same arithmetic, a hundredth of the time
The clearest way to feel the difference is to price a symmetrical exchange: two
thousand words in, two thousand words out.
The lesson stops here
5 more paragraphs to go
You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.
The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.
See the planThe contentsThis is the reading half
Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.
The contents