Serving a Model to Many People
The Average Minute Does Not Exist
Last timeWhere a Request's Time Goes
Requests arrive in clumps, and the clumping is what builds queues. A service sized for its average traffic will be overwhelmed by traffic whose average it was sized for.
A service that can handle two hundred requests per second, receiving one hundred
requests per second, sounds like a service with nothing to worry about. It has
twice the capacity it needs. Yet its users wait, and during some minutes they
wait a long time. The reason is not a bug and not a shortage of hardware. It is
that the hundred requests do not arrive one every ten milliseconds.
Why there is a queue when there is room
Capacity is perishable. A second in which no request arrived is not banked
anywhere. If the hundred requests per second arrive as nothing for half a second
and then a hundred in the other half, the service spends half its time idle and
half its time two hundred per second deep, and the requests in the second half
wait. Their waiting is not cancelled out by the idleness that preceded it.
That is the whole mechanism, and it does not require anything unusual. Arrivals
that are merely independent of one another already clump, because independence
means nothing coordinates them to keep their distance.
- the average time a request spends waiting before anyone starts on it
- the unevenness factor, which is one when arrivals are perfectly random and larger when they clump
- the spread of the gaps between arrivals, measured relative to the average gap
- the load factor, which is how busy the service is divided by how idle it is
- the time it takes to serve one request once started
What unevenness costs
Read the first factor carefully, because it is where the surprises live. Traffic
whose gaps are all identical has a spread of zero and the factor is a half.
Traffic that is merely random has a spread of one and the factor is one. Traffic
that arrives in clumps, which is what retries and scheduled jobs produce, can
easily have a spread of three, and the factor is then five.
So two services with the same average arrival rate, the same hardware and the
same service time can differ tenfold in what their users experience, with the
entire difference sitting in how the work showed up.
The lesson stops here
4 more paragraphs to go
You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.
The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.
See the planThe contentsThis is the reading half
Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.
The contents