ContentsThe library

Serving a Model to Many People

A Number You Can Be Held To

Last timeCapacity Arrives Late

A latency promise is only meaningful if it names a percentile, a window, a place of measurement and what counts as a failure. Then it becomes a design constraint rather than an aspiration.

Most latency promises cannot be broken, because they cannot be checked. "Under

two seconds" is a sentence two people can read, measure honestly, and disagree

about by a factor of ten. One of them took the median at the server, the other

took the ninety-ninth percentile at the phone, and both of them believe they

have answered the question.

A promise worth writing has four parts, and omitting any one of them is what

produces the argument.

FIG 1Four ways of writing the same intention, judged on four things
two people compute the scan be reported automaticonstrains the design, 1survives an argument, 1
fast responses1111
under two seconds2212
99th percentile under tw3333
99th percentile of first4444
The last row is tedious to read and is the only one that can be kept or broken. Every missing part is a place where two honest measurements diverge, and the divergence is always discovered during an incident rather than before one.

Whose clock

Of the four parts, the place of measurement is the one most often got wrong,

and it is wrong in a consistent direction.

A server's timer starts when the request arrives at the server. That excludes

the name lookup, the connection, the time spent in a queue in front of the load

balancer, and the network in both directions. Those stages are not small and,

more importantly, they are where the worst cases live. A request that was

dropped before it reached the service is counted by nobody at all, which means

the server's own numbers improve as the service gets worse.

So the promise should be measured where the person waiting is standing. In

practice that means instrumenting the client, or running a probe that makes real

requests from outside, and treating the server's own timings as a diagnostic

rather than as the promise.

FIG 2Share of requests inside a budget, for three services
0.000.300.600.901.201.025.850.575.3100.0budget, in tens of milliseconds
service Aservice Bservice C
A promise is a point on one of these curves. Service A reaches ninety-nine percent at about nine hundred milliseconds and can promise that. Service B never reaches it within the range shown, so any ninety-ninth percentile promise it makes under a second is a promise it will break most weeks.

The allowance hidden inside the promise

A promise of ninety-nine percent inside the threshold says something else at the

same time, which teams routinely fail to notice. It says that one percent

outside the threshold is permitted.

The lesson stops here

5 more paragraphs to go

You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.

The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.

See the planThe contents

This is the reading half

Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.

The contents

The rest of this course

  1. 01Mostly Waiting
  2. 02The Average Minute Does Not Existopening only
  3. 03Reading the Weights Once for Everybodyopening only
  4. 04The Two Jobs That Get in Each Other's Wayopening only
  5. 05Where the Complaints Actually Liveopening only
  6. 06Saying No While You Still Canopening only
  7. 07Help That Arrives Nine Minutes Lateopening only
  8. 08A Number You Can Be Held Toyou are here

Read alongside