The library

The systems around the model

Be able to explain why a system spread across several machines behaves differently in kind rather than in degree, what copies of data cost, what a network partition forces you to choose, how machines agree on anything at all, and which of these problems you can avoid rather than solve

When One Machine Is Not Enough

One machine either works or it does not. Several machines can be partly working, disagree about what happened, and all be telling the truth.

8 lessons, written and corrected before you arrived. Reading them here needs no account. The first reads the whole way through; the others open and then stop, because a page nobody owns cannot tell who is reading it. Starting the course gives you your own copy, where every idea has problems standing under it and you can ask about any sentence.

Start reading

  1. 01The Three Reasons, and What Each One CostsA system spreads across machines for capacity, for survival, or for distance, and each of those buys something specific at the price of a failure mode that did not exist before.
  2. 02How Far Behind a Copy Is Allowed to Beopening onlyEvery copy of the data is behind the others by some amount, and the only real decisions are how far behind is acceptable and who is made to wait while it catches up.
  3. 03Silence Means Nothing In Particularopening onlyA message that has not arrived looks exactly like a machine that has died, and no amount of care distinguishes them, which is why every timeout in a system is a guess.
  4. 04What You Give Up When the Network Splitsopening onlyWhen machines cannot reach each other they must either answer with data that may be wrong or refuse to answer at all, and no cleverness produces a third option.
  5. 05How a Group Decides Without a Bossopening onlySeveral machines can agree on a single value even though some are down and messages are lost, by insisting that nothing counts until a majority has seen it.
  6. 06Which Thing Happened Firstopening onlySeparate machines cannot agree on what time it is, so ordering events by their timestamps quietly produces wrong answers, and the fix is to count rather than to read a clock.
  7. 07Cutting the Data Into Piecesopening onlyWhen the data no longer fits on one machine it has to be cut up, and almost every difficulty that follows comes from the choice of where to cut.
  8. 08Designing for the Failure You Will Actually Getopening onlyIn a fleet of any size something is always broken, so the useful question is not how to prevent failure but which parts of the system are allowed to fail without taking the rest with them.