ContentsThe library

Serving a Model to Many People

The Average Minute Does Not Exist

Last timeWhere a Request's Time Goes

Requests arrive in clumps, and the clumping is what builds queues. A service sized for its average traffic will be overwhelmed by traffic whose average it was sized for.

A service that can handle two hundred requests per second, receiving one hundred

requests per second, sounds like a service with nothing to worry about. It has

twice the capacity it needs. Yet its users wait, and during some minutes they

wait a long time. The reason is not a bug and not a shortage of hardware. It is

that the hundred requests do not arrive one every ten milliseconds.

FIG 1Where a burst comes from
Only the first source is inherent. The other four are design decisions, which means most of the burstiness a service experiences is burstiness somebody built, and can therefore be unbuilt by spreading schedules, jittering retries and staggering expiry times.

Why there is a queue when there is room

Capacity is perishable. A second in which no request arrived is not banked

anywhere. If the hundred requests per second arrive as nothing for half a second

and then a hundred in the other half, the service spends half its time idle and

half its time two hundred per second deep, and the requests in the second half

wait. Their waiting is not cancelled out by the idleness that preceded it.

That is the whole mechanism, and it does not require anything unusual. Arrivals

that are merely independent of one another already clump, because independence

means nothing coordinates them to keep their distance.

FIG 2The wait, with unevenness included
the average time a request spends waiting before anyone starts on it
the unevenness factor, which is one when arrivals are perfectly random and larger when they clump
the spread of the gaps between arrivals, measured relative to the average gap
the load factor, which is how busy the service is divided by how idle it is
the time it takes to serve one request once started
Three independent things multiply together, and only the middle one appears in most capacity conversations. Halving the unevenness halves the wait exactly as surely as reducing the load does, and it is often much cheaper to arrange.

What unevenness costs

Read the first factor carefully, because it is where the surprises live. Traffic

whose gaps are all identical has a spread of zero and the factor is a half.

Traffic that is merely random has a spread of one and the factor is one. Traffic

that arrives in clumps, which is what retries and scheduled jobs produce, can

easily have a spread of three, and the factor is then five.

So two services with the same average arrival rate, the same hardware and the

same service time can differ tenfold in what their users experience, with the

entire difference sitting in how the work showed up.

The lesson stops here

4 more paragraphs to go

You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.

The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.

See the planThe contents

This is the reading half

Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.

The contents

The rest of this course

  1. 01Mostly Waiting
  2. 02The Average Minute Does Not Existyou are here
  3. 03Reading the Weights Once for Everybodyopening only
  4. 04The Two Jobs That Get in Each Other's Wayopening only
  5. 05Where the Complaints Actually Liveopening only
  6. 06Saying No While You Still Canopening only
  7. 07Help That Arrives Nine Minutes Lateopening only
  8. 08A Number You Can Be Held Toopening only

Read alongside