ContentsThe library

Serving a Model to Many People

Where the Complaints Actually Live

Last timeReading and Writing Compete

Nobody experiences an average. They experience their own request, and if their page needs twenty of them, they experience the slowest of twenty rather than the typical one.

A service reports an average latency of two hundred milliseconds. Somebody

complains that it took four seconds. Both statements are true, and the average

was never going to reveal the second one.

Latency distributions are not symmetric. Most requests cluster near a floor set

by the actual work, and a thin tail stretches far to the right, made of requests

that were briefly unlucky. The mean of that shape sits above almost everything

and below the part people notice.

FIG 1Four services with similar averages and very different tails
median, ms95th percentile, ms99th percentile, ms99.9th percentile, ms
service A180210260310
service B1402209002400
service C951904204800
service D205230250270
Service C has the best median and the worst catastrophe, and an average would have ranked it first. The marked cell is the number its users will describe to each other, and it is twenty-five times its own typical case.

Which bad case matters to you

The percentile worth promising is not a matter of taste. It is set by how many

requests one user action needs.

If opening a page makes a single request, the ninety-fifth percentile describes

something one user in twenty notices once. If opening a page makes a hundred

requests, which is entirely ordinary for an assembled page, then almost every

page load contains a request from the ninety-ninth percentile. That figure is no

longer the rare case. It is the typical experience.

This is also why a promise written in terms of an average is close to

unenforceable. An average can be met while a tenth of the traffic is terrible,

and it can be missed by a single genuinely broken minute in an otherwise healthy

day. A percentile does neither, because it counts requests rather than

weighing them.

FIG 2The chance of at least one slow request
the chance that the user action is slow, because at least one of its requests was
the chance that any one request is slow
how many requests the action makes
the chance that every single one of them was fine
The whole difficulty is in the exponent. Nothing about the service has to change for a rare event to become a common one; the page simply has to make more requests, which it does every time a feature is added.
FIG 3How often a user action is slow, against how many requests it makes
0.000.300.600.901.201.025.850.575.3100.0requests made by one user action
one request in twenty is slowone in a hundredone in a thousand
Read the middle curve at twenty requests and it is already eighteen percent. Making a per-request slow case ten times rarer buys about one order of magnitude of fan-out, which is one or two years of ordinary feature work.
FIG 4One page load, which needs twenty requests
steprequests completedslowest so far, mschance a slow one has appeared, percentpage rendered yetwhat happened
1116010The first request returns quickly, as almost all of them do.
21224011.40Twelve back, all healthy. Nothing so far suggests this page will be slow.
319210017.40The nineteenth was unlucky and took two seconds. The page is now waiting on it alone.
420210018.21The page renders after two point one seconds, which is the slowest request rather than the typical one.
4 steps
Nineteen of the twenty requests were fine and the page was slow anyway. This is why reducing the median of a fan-out service is almost worthless and reducing its ninety-ninth percentile is almost everything.

Where the slow ones come from

The instinct is to look for a fault, and there usually is not one. The common

causes are shared and momentary.

Arrivals clumped for a second and this request landed in the clump. A different

tenant on the same machine was briefly using the memory bandwidth. A maintenance

task ran. Somebody's two-thousand-word prompt was admitted into the group just

ahead. A retry storm somewhere else raised the load for three seconds.

Every one of these is transient, shared, and invisible to a measurement of any

single component. None of them can be reproduced on demand, which is why tail

investigations run for weeks and conclude that nothing is wrong.

The practical consequence is a change of goal. Rather than hunting for the cause

of a four-second response, which is usually a combination that existed for one

second and will never recur, the work is to make the service tolerate such

moments. That is a different kind of engineering, and it is mostly about not

making any single request depend on any single momentary condition.

The lesson stops here

3 more paragraphs to go

You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.

The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.

See the planThe contents

This is the reading half

Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.

The contents

The rest of this course

  1. 01Mostly Waiting
  2. 02The Average Minute Does Not Existopening only
  3. 03Reading the Weights Once for Everybodyopening only
  4. 04The Two Jobs That Get in Each Other's Wayopening only
  5. 05Where the Complaints Actually Liveyou are here
  6. 06Saying No While You Still Canopening only
  7. 07Help That Arrives Nine Minutes Lateopening only
  8. 08A Number You Can Be Held Toopening only

Read alongside