Serving a Model to Many People
Mostly Waiting
A request's total time is what it spent waiting plus what it spent being served, and under load the first term is several times the second while the hardware looks fine.
A model that takes forty milliseconds to produce a word takes forty milliseconds
to produce a word whether one person is asking or ten thousand are. That number
is a property of the model and the memory it lives in, and almost nothing you do
to the service changes it.
So when a service that was comfortable at lunchtime is unusable at eight in the
evening, and every dashboard says the hardware is fine, the extra time has to
have come from somewhere else. It came from waiting.
The two halves
The serving half is the honest, predictable part. Reading the prompt takes an
amount of time proportional to its length. Producing each word takes an amount of
time set by how fast the model's numbers can be read out of memory. Both are
measurable once and stable.
The waiting half is not stable at all. It depends on how many other requests are
in front of this one, which depends on how close the arrival rate has come to the
service's capacity.
- the total time from arrival to completion, which is what the user experiences
- the time the request actually takes to serve, which is a property of the model and the machine
- the share of the time the service is busy, between nothing and one
- the multiplier that busyness applies to the serving time, which is one at an idle service and ten at ninety percent busy
The two clocks
A generated answer is not one event, it is a first word followed by a stream of
further words, and people react to those separately.
The time until the first word appears is dominated by the queue and by how long
the prompt took to read. The rate at which words arrive afterwards is dominated
by the model and by how many other requests are being served alongside this one.
Those two numbers move independently, and a service can have an excellent average
response time while every single user is staring at an empty box for four
seconds.
| step | milliseconds since the request arrived | requests ahead of this one | words produced so far | milliseconds of actual computation so far | what happened |
|---|---|---|---|---|---|
| 1 | 0 | 23 | 0 | 0 | Arrival. Twenty-three requests are already waiting, and nothing will happen to this one for a while. |
| 2 | 820 | 7 | 0 | 0 | Still waiting. Eight hundred milliseconds have passed and the fourth column has not moved at all. |
| 3 | 1180 | 0 | 0 | 90 | The prompt has been read. The user has now waited nearly one and two tenths of a second to see nothing. |
| 4 | 1220 | 0 | 1 | 130 | The first word appears. This is the moment the user judges the service on. |
| 5 | 4800 | 0 | 92 | 3710 | The answer completes. Of four and eight tenths of a second, about three quarters was computation and the first quarter was queue. |
Counting what is inside
One relation makes queue measurements immediately useful. The number of requests
inside the service at any moment equals the rate at which they arrive multiplied
by the average time each one spends inside.
- the average number of requests inside the service at any instant, which includes both those waiting and those being served
- the rate at which requests arrive, in requests per second
- the average total time a request spends inside, in seconds
The same relation runs the other way and is how capacity gets planned. If you
intend to hold total time under half a second at four hundred requests per
second, the service must be able to hold two hundred requests in progress without
its serving time degrading.
Measuring from the wrong place
The last point is a measurement discipline rather than a piece of theory. A
service that starts its timer when it dequeues a request is measuring only the
serving half. That half is the one that does not get worse under load, so the
graph stays flat while the complaints arrive.
The timing that matters is taken where the request entered, at the edge, with the
clock running from the moment the connection was accepted. Everything else is a
component measurement, useful for diagnosis and useless as a promise.
What to hold on to
A request's time is waiting plus serving. Serving is stable and set by the
hardware; waiting is unstable and set by occupancy, and the multiplier it applies
is one over one minus the busy share. Measure two clocks rather than one, since
the pause before the first word and the rate of words afterwards have different
causes. Use the relation between queue depth, arrival rate and total time to turn
one into another. And take the measurement at the edge, because the service's own
view omits precisely the term that moves.
Recap
- Total time splits into waiting and serving, and these two terms behave completely differently. Serving time is a property of the model and the machine. Waiting time is a property of how busy the machine is, and it grows without bound well before the machine is full.
- A model request has two clocks, the time until the first word appears and the rate at which words arrive afterwards, and a single average latency figure describes neither of them usefully.
- The number of requests inside the service at any moment equals the arrival rate multiplied by the average time each one spends there, which makes queue depth a direct measurement of latency rather than a separate thing to watch.
This is the reading half
Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.
The contents