ContentsThe library

Serving a Model to Many People

Mostly Waiting

A request's total time is what it spent waiting plus what it spent being served, and under load the first term is several times the second while the hardware looks fine.

A model that takes forty milliseconds to produce a word takes forty milliseconds

to produce a word whether one person is asking or ten thousand are. That number

is a property of the model and the memory it lives in, and almost nothing you do

to the service changes it.

So when a service that was comfortable at lunchtime is unusable at eight in the

evening, and every dashboard says the hardware is fine, the extra time has to

have come from somewhere else. It came from waiting.

FIG 1What a request passes through, and where its time accumulates
Two of the first three boxes are pure waiting. The common mistake is to tune the fourth and fifth boxes, which are governed by hardware, when the complaint is about the second, which is governed by arrival rate.

The two halves

The serving half is the honest, predictable part. Reading the prompt takes an

amount of time proportional to its length. Producing each word takes an amount of

time set by how fast the model's numbers can be read out of memory. Both are

measurable once and stable.

The waiting half is not stable at all. It depends on how many other requests are

in front of this one, which depends on how close the arrival rate has come to the

service's capacity.

FIG 2Total time as a function of how busy the service is
the total time from arrival to completion, which is what the user experiences
the time the request actually takes to serve, which is a property of the model and the machine
the share of the time the service is busy, between nothing and one
the multiplier that busyness applies to the serving time, which is one at an idle service and ten at ninety percent busy
Read the multiplier rather than the formula. Going from fifty percent busy to ninety percent busy is a fivefold increase in what users experience, bought with no change to any hardware, and it is the single most common cause of a service that got slower without anything being deployed.
FIG 3Total time against occupancy, for two different service times
0.001250.002500.003750.005000.000.00.20.50.70.9share of the time the service is busy
a fifty millisecond requesta two hundred millisecond request
Both curves are flat for most of their range and then turn vertical, and the turn is in the same place for both. That is the practical lesson: the thing that determines whether you are on the flat part or the steep part is occupancy, and it is not the same thing as whether the machine looks busy to an operating system.

The two clocks

A generated answer is not one event, it is a first word followed by a stream of

further words, and people react to those separately.

The time until the first word appears is dominated by the queue and by how long

the prompt took to read. The rate at which words arrive afterwards is dominated

by the model and by how many other requests are being served alongside this one.

Those two numbers move independently, and a service can have an excellent average

response time while every single user is staring at an empty box for four

seconds.

FIG 4One request's life inside a busy service, measured in milliseconds
stepmilliseconds since the request arrivedrequests ahead of this onewords produced so farmilliseconds of actual computation so farwhat happened
102300Arrival. Twenty-three requests are already waiting, and nothing will happen to this one for a while.
2820700Still waiting. Eight hundred milliseconds have passed and the fourth column has not moved at all.
311800090The prompt has been read. The user has now waited nearly one and two tenths of a second to see nothing.
4122001130The first word appears. This is the moment the user judges the service on.
548000923710The answer completes. Of four and eight tenths of a second, about three quarters was computation and the first quarter was queue.
5 steps
Compare the first and last columns at each step. The gap between them is pure waiting, and it is all concentrated before the user sees anything, which is the worst place for it to be.

Counting what is inside

One relation makes queue measurements immediately useful. The number of requests

inside the service at any moment equals the rate at which they arrive multiplied

by the average time each one spends inside.

FIG 5The relation between queue depth, arrival rate and latency
the average number of requests inside the service at any instant, which includes both those waiting and those being served
the rate at which requests arrive, in requests per second
the average total time a request spends inside, in seconds
This holds whatever the arrival pattern and whatever order the queue is served in, which is rare for a result this useful. In practice it means a queue depth gauge is a latency gauge: at two hundred arrivals per second, a steady depth of sixty requests means each one is spending three tenths of a second inside.

The same relation runs the other way and is how capacity gets planned. If you

intend to hold total time under half a second at four hundred requests per

second, the service must be able to hold two hundred requests in progress without

its serving time degrading.

Measuring from the wrong place

The last point is a measurement discipline rather than a piece of theory. A

service that starts its timer when it dequeues a request is measuring only the

serving half. That half is the one that does not get worse under load, so the

graph stays flat while the complaints arrive.

The timing that matters is taken where the request entered, at the edge, with the

clock running from the moment the connection was accepted. Everything else is a

component measurement, useful for diagnosis and useless as a promise.

What to hold on to

A request's time is waiting plus serving. Serving is stable and set by the

hardware; waiting is unstable and set by occupancy, and the multiplier it applies

is one over one minus the busy share. Measure two clocks rather than one, since

the pause before the first word and the rate of words afterwards have different

causes. Use the relation between queue depth, arrival rate and total time to turn

one into another. And take the measurement at the edge, because the service's own

view omits precisely the term that moves.

Recap

  • Total time splits into waiting and serving, and these two terms behave completely differently. Serving time is a property of the model and the machine. Waiting time is a property of how busy the machine is, and it grows without bound well before the machine is full.
  • A model request has two clocks, the time until the first word appears and the rate at which words arrive afterwards, and a single average latency figure describes neither of them usefully.
  • The number of requests inside the service at any moment equals the arrival rate multiplied by the average time each one spends there, which makes queue depth a direct measurement of latency rather than a separate thing to watch.

This is the reading half

Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.

The contents

NextTraffic Does Not Arrive Evenly →

The rest of this course

  1. 01Mostly Waitingyou are here
  2. 02The Average Minute Does Not Existopening only
  3. 03Reading the Weights Once for Everybodyopening only
  4. 04The Two Jobs That Get in Each Other's Wayopening only
  5. 05Where the Complaints Actually Liveopening only
  6. 06Saying No While You Still Canopening only
  7. 07Help That Arrives Nine Minutes Lateopening only
  8. 08A Number You Can Be Held Toopening only

Read alongside