ContentsThe library

Serving a Model to Many People

Help That Arrives Nine Minutes Late

Last timeTurning Work Away on Purpose

Adding a machine is a decision whose effect appears minutes later. Everything about automatic scaling follows from that delay, including why spare capacity has to be paid for.

An automatic scaler is usually described as though it closes a loop. Demand

rises, the scaler notices, capacity appears, the service copes. The description

is accurate except for one word, which is "appears". Capacity does not appear.

It is requested, and it arrives several minutes later, and almost everything

difficult about the subject follows from that gap.

FIG 1The chain between a decision and a served request
Each stage looks modest and the sum does not. Five minutes is a good outcome, fifteen is common, and the figure worth knowing for your own service is the measured time from threshold to first served request rather than any single stage.

The signal has to lead, not lag

Once you accept the delay, the choice of scaling signal stops being a matter of

taste. A signal that only moves after the service is in trouble buys nothing,

because the help it summons will arrive several minutes after that.

Occupancy is the common choice and the weakest one. Its trouble is that it

saturates. A service that is comfortably busy reads ninety-five percent, and a

service with a two-minute backlog also reads ninety-five percent, so the number

stops distinguishing between the two cases exactly when the distinction is the

whole question.

Queue depth does not saturate. Neither does the age of the oldest request still

waiting. Both keep rising for as long as the arrivals exceed the service rate,

which makes them honest in overload and, more usefully, early. A queue that is

two seconds deep and growing is a reliable statement about the next few minutes,

and the next few minutes is precisely the horizon the scaler has to act on.

FIG 2Demand and the capacity chasing it
0.0050.00100.00150.00200.001.015.830.545.360.0minute
demandcapacity, three-minute delaycapacity, nine-minute delay
A lagging capacity line is short of demand on every rise and in surplus on every fall, so the service is both slow and expensive. The nine-minute line spends part of each cycle in opposition to the demand it is meant to be following.

Headroom is what the delay costs

Here is the consequence that people resist. During the delay, no new capacity

exists. Every additional request that arrives in those minutes must be absorbed

by machines that are already running. So the spare capacity you carry is not

slack to be eliminated, it is the only thing standing between a rise in demand

and a queue.

The lesson stops here

4 more paragraphs to go

You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.

The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.

See the planThe contents

This is the reading half

Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.

The contents

The rest of this course

  1. 01Mostly Waiting
  2. 02The Average Minute Does Not Existopening only
  3. 03Reading the Weights Once for Everybodyopening only
  4. 04The Two Jobs That Get in Each Other's Wayopening only
  5. 05Where the Complaints Actually Liveopening only
  6. 06Saying No While You Still Canopening only
  7. 07Help That Arrives Nine Minutes Lateyou are here
  8. 08A Number You Can Be Held Toopening only

Read alongside