ContentsThe library

When One Machine Is Not Enough

The Three Reasons, and What Each One Costs

A system spreads across machines for capacity, for survival, or for distance, and each of those buys something specific at the price of a failure mode that did not exist before.

Before anything else in this course, a question worth asking out loud: does this

system need more than one machine at all? The answer is yes more often than not,

but the reason matters enormously, and systems built without naming the reason

tend to pay for all three and receive one.

One machine is bigger than you think

A server you can rent today has dozens of cores and a terabyte of memory. A

carefully written single process on that machine will serve tens of thousands of

requests a second against data that fits in its memory, with no network hop

anywhere in the path and no possibility of two parts of the system disagreeing.

FIG 1What more machines actually return
0.0010.0020.0030.0040.001.08.816.524.332.0machines
what people assume: each machine adds its full sharewhat is usually measured
The gap between the two lines is coordination. The share of the work that has to be agreed between machines does not get faster when machines are added, so it eventually dominates, and the curve flattens well before the fleet does.

The flattening is not a defect of anybody's implementation. It follows from the

fact that some fraction of the work has to be agreed between machines, and that

fraction costs a network round trip rather than a function call. A round trip

inside one data centre is around half a millisecond; a function call is a few

nanoseconds. A hundred thousand times slower, on the part of the work that cannot

be parallelised away.

FIG 2Three reasons, three different arrangements
Naming which reason applies determines the arrangement. A system built for capacity and then asked to survive a failure usually has to be rebuilt rather than extended.

The three reasons

It is worth measuring before accepting the conclusion. A workload that needs a

fleet when written one way frequently fits on a single machine when written

another, because the cost of crossing the network dwarfs almost every

inefficiency inside one process.

It is usually survival

Capacity is the reason people talk about and survival is the reason most systems

are actually spread out. The arithmetic is unforgiving: a machine that fails once

every three years is extremely reliable, and in a fleet of a thousand such

machines something fails every day. The data on a failed machine is gone unless it

exists somewhere else.

FIG 3What copies buy, if failures are independent
the chance the system as a whole is available
the chance any single machine is available
how many machines hold a copy
the chance one machine is down
the chance every copy is down at the same moment, which is the only way the system is lost
The exponent is what makes copies worth their cost. Going from one copy to three turns a one percent chance of being down into a one in a million chance, provided the failures really are independent, which is the assumption that does most of the lying in practice.
FIG 4Adding copies, with a one percent chance each is down
stepmachines holding a copychance all are down, in parts per milliondowntime minutes per yearmonthly cost in thousandswhat happened
111e+0452564One machine. Nearly four days of downtime a year, which is what a single machine actually gives you.
2210052.68A second copy. Two orders of magnitude better for twice the money, which is the best trade in the whole table.
3310.512A third. Another two orders of magnitude, and now half a minute a year.
450.00015e-0520Five copies. The arithmetic says essentially never, and at this point the arithmetic has stopped being the limiting factor.
4 steps
The last row is the one to be suspicious of. Beyond three copies the calculated availability is far better than anything the surrounding reality supports, because the failures are no longer independent: a bad deployment, a configuration error or a power event takes copies down together.

Independence is also something you can buy, by placing copies in different racks,

different buildings and different regions, and each step of that buys less

correlation at the price of distance. The three reasons meet here, which is why

the arrangement question is rarely answered once.

What you buy along with it

Here is the part that is not on the invoice. With one machine there are two

states, working and not, and the machine knows which one it is in. With two

machines there is a third.

A request arrives at the second machine, is carried out, and the acknowledgement

is lost on the way back. The first machine cannot tell this apart from the request

never arriving. The second machine is alive and healthy and unreachable, which it

cannot tell apart from being isolated, and the first machine cannot tell apart

from the second having died. Nobody has made a mistake and the system is now in a

state that has no equivalent on a single machine.

FIG 5What each arrangement actually provides
capacity for writessurvives losing one machsimplicityread speed nearby
one machine0.200.101.000.30
copies for reading0.250.900.800.90
copies for survival0.300.950.550.40
data split into pieces0.950.250.450.35
The marked cell is the one that gets given up first and missed longest. A single machine is the only arrangement where the state of the system is knowable, and every column gained after that is paid for in the third one.

What to hold on to

Ask which of the three reasons applies before choosing an arrangement, because

dividing data and copying data are opposite moves and only one of them solves

each problem. Expect the second machine to buy a great deal and the twentieth to

buy very little, since the coordinated share of the work sets a ceiling. And

expect the real cost to be the new middle state, where a machine is neither

working nor failed and nothing in the system can tell which.

Recap

  • One machine is far larger than most people assume, so the honest first question is whether the workload has actually outgrown it or merely been written as though it had.
  • The three reasons to add machines are capacity, survival and distance, and they want different arrangements, so starting by naming which one you are buying prevents most of the expensive mistakes.
  • The thing you buy alongside all three is partial failure: a state in which the system is neither up nor down, which no amount of care on a single machine ever produces.

This is the reading half

Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.

The contents

NextKeeping Copies In Step →

The rest of this course

  1. 01The Three Reasons, and What Each One Costsyou are here
  2. 02How Far Behind a Copy Is Allowed to Beopening only
  3. 03Silence Means Nothing In Particularopening only
  4. 04What You Give Up When the Network Splitsopening only
  5. 05How a Group Decides Without a Bossopening only
  6. 06Which Thing Happened Firstopening only
  7. 07Cutting the Data Into Piecesopening only
  8. 08Designing for the Failure You Will Actually Getopening only

Read alongside