When One Machine Is Not Enough
The Three Reasons, and What Each One Costs
A system spreads across machines for capacity, for survival, or for distance, and each of those buys something specific at the price of a failure mode that did not exist before.
Before anything else in this course, a question worth asking out loud: does this
system need more than one machine at all? The answer is yes more often than not,
but the reason matters enormously, and systems built without naming the reason
tend to pay for all three and receive one.
One machine is bigger than you think
A server you can rent today has dozens of cores and a terabyte of memory. A
carefully written single process on that machine will serve tens of thousands of
requests a second against data that fits in its memory, with no network hop
anywhere in the path and no possibility of two parts of the system disagreeing.
The flattening is not a defect of anybody's implementation. It follows from the
fact that some fraction of the work has to be agreed between machines, and that
fraction costs a network round trip rather than a function call. A round trip
inside one data centre is around half a millisecond; a function call is a few
nanoseconds. A hundred thousand times slower, on the part of the work that cannot
be parallelised away.
The three reasons
It is worth measuring before accepting the conclusion. A workload that needs a
fleet when written one way frequently fits on a single machine when written
another, because the cost of crossing the network dwarfs almost every
inefficiency inside one process.
It is usually survival
Capacity is the reason people talk about and survival is the reason most systems
are actually spread out. The arithmetic is unforgiving: a machine that fails once
every three years is extremely reliable, and in a fleet of a thousand such
machines something fails every day. The data on a failed machine is gone unless it
exists somewhere else.
- the chance the system as a whole is available
- the chance any single machine is available
- how many machines hold a copy
- the chance one machine is down
- the chance every copy is down at the same moment, which is the only way the system is lost
| step | machines holding a copy | chance all are down, in parts per million | downtime minutes per year | monthly cost in thousands | what happened |
|---|---|---|---|---|---|
| 1 | 1 | 1e+04 | 5256 | 4 | One machine. Nearly four days of downtime a year, which is what a single machine actually gives you. |
| 2 | 2 | 100 | 52.6 | 8 | A second copy. Two orders of magnitude better for twice the money, which is the best trade in the whole table. |
| 3 | 3 | 1 | 0.5 | 12 | A third. Another two orders of magnitude, and now half a minute a year. |
| 4 | 5 | 0.0001 | 5e-05 | 20 | Five copies. The arithmetic says essentially never, and at this point the arithmetic has stopped being the limiting factor. |
Independence is also something you can buy, by placing copies in different racks,
different buildings and different regions, and each step of that buys less
correlation at the price of distance. The three reasons meet here, which is why
the arrangement question is rarely answered once.
What you buy along with it
Here is the part that is not on the invoice. With one machine there are two
states, working and not, and the machine knows which one it is in. With two
machines there is a third.
A request arrives at the second machine, is carried out, and the acknowledgement
is lost on the way back. The first machine cannot tell this apart from the request
never arriving. The second machine is alive and healthy and unreachable, which it
cannot tell apart from being isolated, and the first machine cannot tell apart
from the second having died. Nobody has made a mistake and the system is now in a
state that has no equivalent on a single machine.
| capacity for writes | survives losing one mach | simplicity | read speed nearby | |
|---|---|---|---|---|
| one machine | 0.20 | 0.10 | 1.00 | 0.30 |
| copies for reading | 0.25 | 0.90 | 0.80 | 0.90 |
| copies for survival | 0.30 | 0.95 | 0.55 | 0.40 |
| data split into pieces | 0.95 | 0.25 | 0.45 | 0.35 |
What to hold on to
Ask which of the three reasons applies before choosing an arrangement, because
dividing data and copying data are opposite moves and only one of them solves
each problem. Expect the second machine to buy a great deal and the twentieth to
buy very little, since the coordinated share of the work sets a ceiling. And
expect the real cost to be the new middle state, where a machine is neither
working nor failed and nothing in the system can tell which.
Recap
- One machine is far larger than most people assume, so the honest first question is whether the workload has actually outgrown it or merely been written as though it had.
- The three reasons to add machines are capacity, survival and distance, and they want different arrangements, so starting by naming which one you are buying prevents most of the expensive mistakes.
- The thing you buy alongside all three is partial failure: a state in which the system is neither up nor down, which no amount of care on a single machine ever produces.
This is the reading half
Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.
The contents