When One Machine Is Not Enough
Designing for the Failure You Will Actually Get
Last timeCutting the Data Into Pieces
In a fleet of any size something is always broken, so the useful question is not how to prevent failure but which parts of the system are allowed to fail without taking the rest with them.
A single machine is either working or not, and that binary is comfortable. A
thousand machines are never all working. Some fraction of disks fails each year,
machines reboot for updates, and networks lose packets continuously. At fleet
scale all of those become rates rather than events.
| step | machines in the fleet | disks failing per year | machines down right now | how often this is remarkable | what happened |
|---|---|---|---|---|---|
| 1 | 1 | 0.03 | 0.001 | 1 | One machine. A disk failure is a memorable event that happens roughly once every thirty years. |
| 2 | 100 | 3 | 0.1 | 0.4 | A hundred machines. Three disk replacements a year, which is a ticket rather than an incident. |
| 3 | 1000 | 30 | 1.4 | 0 | A thousand. Something is always down, and a disk is replaced most weeks. Nobody is told. |
| 4 | 1e+04 | 300 | 14 | 0 | Ten thousand. Fourteen machines are broken at any moment and the system is considered healthy. |
The error handling is where the damage comes from
Here is the most useful empirical finding in this subject. When a large number of
serious failures in widely used systems were examined, most were not caused by
the original fault. They were caused by the code that was supposed to handle it.
| how often it happens | caught by ordinary tests | caught in review | damage if unhandled | |
|---|---|---|---|---|
| a dependency returns an | 0.95 | 0.90 | 0.85 | 0.30 |
| a dependency times out | 0.70 | 0.40 | 0.35 | 0.70 |
| a dependency answers ver | 0.55 | 0.15 | 0.20 | 0.95 |
| a dependency answers wro | 0.40 | 0.05 | 0.10 | 0.95 |
Every mandatory dependency costs you
A service is no more available than the things it insists on. That multiplication
is unforgiving, and it is the reason the single most valuable reliability exercise
is going through the dependency list and asking which entries are genuinely
mandatory.
The lesson stops here
4 more paragraphs to go
You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.
The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.
See the planThe contentsThis is the reading half
Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.
The contents