ContentsThe library

When One Machine Is Not Enough

Designing for the Failure You Will Actually Get

Last timeCutting the Data Into Pieces

In a fleet of any size something is always broken, so the useful question is not how to prevent failure but which parts of the system are allowed to fail without taking the rest with them.

A single machine is either working or not, and that binary is comfortable. A

thousand machines are never all working. Some fraction of disks fails each year,

machines reboot for updates, and networks lose packets continuously. At fleet

scale all of those become rates rather than events.

FIG 1What a thousand machines look like on a normal day
stepmachines in the fleetdisks failing per yearmachines down right nowhow often this is remarkablewhat happened
110.030.0011One machine. A disk failure is a memorable event that happens roughly once every thirty years.
210030.10.4A hundred machines. Three disk replacements a year, which is a ticket rather than an incident.
31000301.40A thousand. Something is always down, and a disk is replaced most weeks. Nobody is told.
41e+04300140Ten thousand. Fourteen machines are broken at any moment and the system is considered healthy.
4 steps
The last row is the one to design for. A system that behaves correctly only when every machine is up describes a state it will never actually be in, so that behaviour is untested by definition.

The error handling is where the damage comes from

Here is the most useful empirical finding in this subject. When a large number of

serious failures in widely used systems were examined, most were not caused by

the original fault. They were caused by the code that was supposed to handle it.

FIG 2One slow dependency becoming a total outage
The upper path is the common shape of a large outage: one component degrades and a shared resource in its callers converts that into everything failing. The lower path requires no cleverness, only that somebody decided in advance what the panel should do when it has no data.
FIG 3Which failure modes actually get handled
how often it happenscaught by ordinary testscaught in reviewdamage if unhandled
a dependency returns an 0.950.900.850.30
a dependency times out0.700.400.350.70
a dependency answers ver0.550.150.200.95
a dependency answers wro0.400.050.100.95
The marked cell is the gap that produces outages. Slow answers are common, almost never tested, and the most damaging of the four, because the calling code usually has no deadline at all and simply waits.

Every mandatory dependency costs you

A service is no more available than the things it insists on. That multiplication

is unforgiving, and it is the reason the single most valuable reliability exercise

is going through the dependency list and asking which entries are genuinely

mandatory.

The lesson stops here

4 more paragraphs to go

You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.

The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.

See the planThe contents

This is the reading half

Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.

The contents

The rest of this course

  1. 01The Three Reasons, and What Each One Costs
  2. 02How Far Behind a Copy Is Allowed to Beopening only
  3. 03Silence Means Nothing In Particularopening only
  4. 04What You Give Up When the Network Splitsopening only
  5. 05How a Group Decides Without a Bossopening only
  6. 06Which Thing Happened Firstopening only
  7. 07Cutting the Data Into Piecesopening only
  8. 08Designing for the Failure You Will Actually Getyou are here

Read alongside