Serving a Model to Many People
Where the Complaints Actually Live
Last timeReading and Writing Compete
Nobody experiences an average. They experience their own request, and if their page needs twenty of them, they experience the slowest of twenty rather than the typical one.
A service reports an average latency of two hundred milliseconds. Somebody
complains that it took four seconds. Both statements are true, and the average
was never going to reveal the second one.
Latency distributions are not symmetric. Most requests cluster near a floor set
by the actual work, and a thin tail stretches far to the right, made of requests
that were briefly unlucky. The mean of that shape sits above almost everything
and below the part people notice.
| median, ms | 95th percentile, ms | 99th percentile, ms | 99.9th percentile, ms | |
|---|---|---|---|---|
| service A | 180 | 210 | 260 | 310 |
| service B | 140 | 220 | 900 | 2400 |
| service C | 95 | 190 | 420 | 4800 |
| service D | 205 | 230 | 250 | 270 |
Which bad case matters to you
The percentile worth promising is not a matter of taste. It is set by how many
requests one user action needs.
If opening a page makes a single request, the ninety-fifth percentile describes
something one user in twenty notices once. If opening a page makes a hundred
requests, which is entirely ordinary for an assembled page, then almost every
page load contains a request from the ninety-ninth percentile. That figure is no
longer the rare case. It is the typical experience.
This is also why a promise written in terms of an average is close to
unenforceable. An average can be met while a tenth of the traffic is terrible,
and it can be missed by a single genuinely broken minute in an otherwise healthy
day. A percentile does neither, because it counts requests rather than
weighing them.
- the chance that the user action is slow, because at least one of its requests was
- the chance that any one request is slow
- how many requests the action makes
- the chance that every single one of them was fine
| step | requests completed | slowest so far, ms | chance a slow one has appeared, percent | page rendered yet | what happened |
|---|---|---|---|---|---|
| 1 | 1 | 160 | 1 | 0 | The first request returns quickly, as almost all of them do. |
| 2 | 12 | 240 | 11.4 | 0 | Twelve back, all healthy. Nothing so far suggests this page will be slow. |
| 3 | 19 | 2100 | 17.4 | 0 | The nineteenth was unlucky and took two seconds. The page is now waiting on it alone. |
| 4 | 20 | 2100 | 18.2 | 1 | The page renders after two point one seconds, which is the slowest request rather than the typical one. |
Where the slow ones come from
The instinct is to look for a fault, and there usually is not one. The common
causes are shared and momentary.
Arrivals clumped for a second and this request landed in the clump. A different
tenant on the same machine was briefly using the memory bandwidth. A maintenance
task ran. Somebody's two-thousand-word prompt was admitted into the group just
ahead. A retry storm somewhere else raised the load for three seconds.
Every one of these is transient, shared, and invisible to a measurement of any
single component. None of them can be reproduced on demand, which is why tail
investigations run for weeks and conclude that nothing is wrong.
The practical consequence is a change of goal. Rather than hunting for the cause
of a four-second response, which is usually a combination that existed for one
second and will never recur, the work is to make the service tolerate such
moments. That is a different kind of engineering, and it is mostly about not
making any single request depend on any single momentary condition.
The lesson stops here
3 more paragraphs to go
You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.
The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.
See the planThe contentsThis is the reading half
Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.
The contents