The Same Thing, Asked Again
Caching works because real traffic is wildly lopsided. A handful of items absorb most of the requests, so a cache far too small to hold everything can still answer nearly everything.
A cache is a bet. You produce an answer once, keep a copy, and bet that somebody
will ask for the same thing before the copy is worthless. Everything technical
about caching follows from whether that bet pays, so it is worth being precise
about why it does.
Nothing is asked for evenly
Imagine a collection of a million items and a million requests arriving for
them. If requests were spread evenly, each item would be asked for once, no item
would ever be asked for twice, and a cache would be useless. That is not what
happens anywhere.
The curve matters more than any single number on it. The most requested item in
a production collection commonly takes a few percent of all traffic by itself.
The tenth takes well under one percent. By the thousandth you are into
requests-per-hour, and by the hundred thousandth into requests-per-week. The
measurements from large production cache fleets show this shape repeatedly, with
the steepness varying by workload but the shape never flattening into the even
line.
Why the first megabyte is worth the most
Because popularity falls away so fast, the relationship between how much you
keep and how much you can answer is not a straight line. Doubling a cache that
already holds the popular items adds very little, because the things it newly
holds are things nobody is asking for.
This has a practical consequence that surprises people who expect caches to be
expensive. You usually do not need a cache sized to your data. You need one
sized to the part of your data that is in demand this hour, which may be a
thousandth of the whole. Arguments about cache capacity are therefore usually
arguments about the wrong thing: the interesting question is almost never how
much memory to buy, it is whether the right things are in the memory you already
have.
What makes one answer worth keeping over another
Popularity alone does not decide it. An answer that takes a second to assemble
and is wanted ten times an hour saves ten seconds an hour. An answer that takes
a tenth of a millisecond and is wanted ten thousand times an hour saves one
second. The expensive, less popular answer is the better thing to keep.
- the work saved per unit of time by keeping this answer
- how often the answer is asked for, per unit of time
- what it costs to produce the answer from scratch
- what it costs to retrieve the kept copy instead
- the saving on a single request, which is what the cache is actually selling
The same arithmetic explains an uncomfortable case. If the kept copy is almost
as expensive to fetch as the original answer is to produce, the difference term
collapses and the cache earns nothing while still costing memory, complexity and
a new way to be wrong. Caches placed across a network from the thing they are
caching sometimes land in exactly this position.
The second request arrives soon
There is a second kind of lopsidedness, and it is in time rather than in
popularity. When an item is asked for, the next request for it tends to arrive
quickly. Production measurements of reuse intervals show most of them in seconds
or minutes, with a long thin tail stretching out to days.
| step | minutes since the item first appeared | requests for it in that minute | distinct other items also in demand | share of requests the cache answered | what happened |
|---|---|---|---|---|---|
| 1 | 0 | 1 | 840 | 0 | The first request. Nothing is kept yet, so this one pays full price. |
| 2 | 3 | 46 | 910 | 0.91 | Something has made the item interesting. Almost every request is now being answered from the copy. |
| 3 | 28 | 12 | 880 | 0.93 | Interest is fading but the copy is still earning. |
| 4 | 240 | 0 | 870 | 0.92 | Nobody wants it any more. The space it occupies is now worth more to something else, which is the subject of a later lesson. |
When the bet does not pay
It is worth knowing the shapes of traffic where none of this applies, because
people do install caches in front of them. If each request carries a unique
search phrase, a one-off identifier, or a freshly assembled personal view, there
is no second request to catch. If the underlying data changes more often than it
is read, the copies are stale before anyone wants them. In both cases the cache
adds memory, a new failure mode and a confusing extra layer, and returns
nothing.
What to hold on to
Caching works because request traffic is steeply uneven in both popularity and
time. A cache that holds a tiny share of your items can answer most of your
requests, the first increment of cache is worth far more than the last, and what
deserves space is set by how often an answer is wanted multiplied by what
producing it costs. Where none of that holds, no cache will help.
Recap
- Requests are not spread evenly over the things that could be requested, and the gap is not small: the most popular item in a collection is often asked for thousands of times more often than the median one, which is the entire reason a small cache can carry a large share of the load.
- What makes an answer worth keeping is the chance it is asked again multiplied by what producing it costs, so a rarely wanted answer that takes a second to build can be a better candidate than a common one that takes a microsecond.
- Reuse is concentrated in time as well as in popularity, so most of what a cache earns comes from requests arriving within minutes of each other rather than from anything held for a long time.
This is the reading half
Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.
The contents