ContentsThe library

Designing an Interface to Call

A Limit Is Only Useful If Somebody Can Obey It

Last timeToo Much to Return at Once

Refusing a request is easy. Designing a limit that a well-meaning caller can actually keep to, and telling them enough to do it, is the part that gets skipped.

What a limit is protecting

There are two reasons to limit a caller and only one of them gets talked

about.

The first is capacity. Your service has a finite amount of work it can do,

and a caller who asks for more than that will get worse service and so will

everybody else, until something falls over. A limit is the mechanism that

turns an overload into a refusal, which is the better of the two outcomes

because it is reversible.

The second reason is fairness, and it is the one that actually matters most

days. Capacity is shared. Without a per-caller limit, one integration that

loops without a delay consumes the whole budget, and every other caller

experiences that as your service being broken. They will report it as your

service being broken, because from where they stand it is. A limit is what

makes one caller's mistake their problem instead of everybody's.

That framing decides something practical: the limit is per caller, not

global. A global cap protects your machines and does nothing about fairness,

so the loudest caller still wins.

FIG 1One caller without a limit, from four points of view
stepminutethat callereverybody elseyour dashboardwhat happened
1140 requests, a normal integrationnormalhealthyNothing to see. This is the state the limit is designed to preserve.
22a retry loop with no delay startsslowerlatency risingA bug on their side, not malice. The most common cause is a retry that forgot to wait.
332,400 requeststiming outsaturatedYour capacity is now spent on one caller. The others are receiving errors for requests that are perfectly reasonable.
442,400 requestsreporting an outageincidentFour separate teams open tickets saying your service is down. It is not down; it is fully occupied by somebody else.
4 steps
The third column is the argument for limits. Nothing in the first column is deliberate and nothing in it is your fault, and the consequence lands entirely on people who did nothing. A limit converts this into one caller receiving refusals.

The shapes a limit can take

Three shapes cover nearly everything, and they differ in what they permit

rather than in how strict they are.

A fixed window counts requests in each clock minute and resets. It is the

simplest to build and it has a specific defect: a caller can spend their

whole allowance at the end of one minute and the whole next allowance at the

start of the following one, putting twice the rate through in two seconds.

The limit was honoured and the burst arrived anyway.

A sliding window counts over the last sixty seconds continuously, which

removes that boundary effect at the cost of remembering individual request

times.

A bucket takes a different view. Allowance accumulates at a steady rate up to

a maximum, each request spends one unit, and a request arriving at an empty

bucket is refused. The maximum is a burst you are deliberately granting, and

the fill rate is the average you are enforcing. That pair is usually what

people mean when they describe what they want, which is why the bucket is the

common choice.

FIG 2What a bucket permits over a period
requests permitted in the period
bucket size, the burst allowed at once
fill rate, requests per second on average
length of the period in seconds
Two numbers describe the whole policy. A bucket of 100 filling at 10 per second lets a caller send 100 immediately and 700 over the following minute. Both halves are deliberate: the burst absorbs normal unevenness, the rate protects the average.
FIG 3One request against a bucket
Nine steps, and the two that are usually missing are the two that report numbers. Enforcement without reporting produces a caller who cannot do anything intelligent except guess.

Telling the caller enough

A refusal that says only that the limit was exceeded is close to useless. The

caller knows they were refused; what they need to know is what to do next.

The lesson stops here

3 more paragraphs to go

You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.

The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.

See the planThe contents

This is the reading half

Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.

The contents

The rest of this course

  1. 01Every Awkward Interface You Have Ever Used Was Awkward for the Same Reason, and the Reason Was Decided in the First Half Hour
  2. 02There Are Only Two Questions Worth Asking About Any Operation, and Neither of Them Is What It Is Calledopening only
  3. 03Your Error Message Is Read by a Program First and a Human Second, and Almost Every Interface Gets That Order Backwardsopening only
  4. 04The Second Call Is Not a Mistakeopening only
  5. 05Counting From the Start Versus Remembering the Placeopening only
  6. 06A Limit Is Only Useful If Somebody Can Obey Ityou are here
  7. 07Hand Back a Receipt, Not an Answeropening only
  8. 08Add, Migrate, Then Removeopening only

Read alongside