ContentsThe library

Keeping the Answer Nearby

The Same Thing, Asked Again

Caching works because real traffic is wildly lopsided. A handful of items absorb most of the requests, so a cache far too small to hold everything can still answer nearly everything.

A cache is a bet. You produce an answer once, keep a copy, and bet that somebody

will ask for the same thing before the copy is worthless. Everything technical

about caching follows from whether that bet pays, so it is worth being precise

about why it does.

Nothing is asked for evenly

Imagine a collection of a million items and a million requests arriving for

them. If requests were spread evenly, each item would be asked for once, no item

would ever be asked for twice, and a cache would be useless. That is not what

happens anywhere.

FIG 1How often each item is asked for, against how popular it is
0.000.300.600.901.201.013.325.537.850.0rank of the item, most popular first
what real traffic looks likewhat an even spread would look like
The flat line is the world in which caching is pointless. The falling curve is the world you actually work in, and the area under its left edge is the whole business case for a cache.

The curve matters more than any single number on it. The most requested item in

a production collection commonly takes a few percent of all traffic by itself.

The tenth takes well under one percent. By the thousandth you are into

requests-per-hour, and by the hundred thousandth into requests-per-week. The

measurements from large production cache fleets show this shape repeatedly, with

the steepness varying by workload but the shape never flattening into the even

line.

FIG 2Where a day of requests went
A thousand items out of a million absorb six hundred thousand of the million requests. Keeping a thousand things is trivially cheap, which is why the first and smallest cache you add is almost always the one that pays best.

Why the first megabyte is worth the most

Because popularity falls away so fast, the relationship between how much you

keep and how much you can answer is not a straight line. Doubling a cache that

already holds the popular items adds very little, because the things it newly

holds are things nobody is asking for.

This has a practical consequence that surprises people who expect caches to be

expensive. You usually do not need a cache sized to your data. You need one

sized to the part of your data that is in demand this hour, which may be a

thousandth of the whole. Arguments about cache capacity are therefore usually

arguments about the wrong thing: the interesting question is almost never how

much memory to buy, it is whether the right things are in the memory you already

have.

What makes one answer worth keeping over another

Popularity alone does not decide it. An answer that takes a second to assemble

and is wanted ten times an hour saves ten seconds an hour. An answer that takes

a tenth of a millisecond and is wanted ten thousand times an hour saves one

second. The expensive, less popular answer is the better thing to keep.

FIG 3What keeping one answer saves you
the work saved per unit of time by keeping this answer
how often the answer is asked for, per unit of time
what it costs to produce the answer from scratch
what it costs to retrieve the kept copy instead
the saving on a single request, which is what the cache is actually selling
Both terms matter and they are independent. This is why a cache in front of a slow report query can be worth more than one in front of a fast lookup that runs a thousand times as often.

The same arithmetic explains an uncomfortable case. If the kept copy is almost

as expensive to fetch as the original answer is to produce, the difference term

collapses and the cache earns nothing while still costing memory, complexity and

a new way to be wrong. Caches placed across a network from the thing they are

caching sometimes land in exactly this position.

The second request arrives soon

There is a second kind of lopsidedness, and it is in time rather than in

popularity. When an item is asked for, the next request for it tends to arrive

quickly. Production measurements of reuse intervals show most of them in seconds

or minutes, with a long thin tail stretching out to days.

FIG 4One item's life in the cache
stepminutes since the item first appearedrequests for it in that minutedistinct other items also in demandshare of requests the cache answeredwhat happened
1018400The first request. Nothing is kept yet, so this one pays full price.
23469100.91Something has made the item interesting. Almost every request is now being answered from the copy.
328128800.93Interest is fading but the copy is still earning.
424008700.92Nobody wants it any more. The space it occupies is now worth more to something else, which is the subject of a later lesson.
4 steps
The cache earned almost everything it was ever going to earn from this item inside half an hour. That is the usual pattern, and it is why forgetting aggressively costs much less than intuition suggests.
FIG 5The two paths a request can take
Notice that the miss path is strictly slower than having no cache at all, since it does the expensive work and then some bookkeeping. A cache is a trade in which misses pay a small tax so that hits can be cheap.

When the bet does not pay

It is worth knowing the shapes of traffic where none of this applies, because

people do install caches in front of them. If each request carries a unique

search phrase, a one-off identifier, or a freshly assembled personal view, there

is no second request to catch. If the underlying data changes more often than it

is read, the copies are stale before anyone wants them. In both cases the cache

adds memory, a new failure mode and a confusing extra layer, and returns

nothing.

What to hold on to

Caching works because request traffic is steeply uneven in both popularity and

time. A cache that holds a tiny share of your items can answer most of your

requests, the first increment of cache is worth far more than the last, and what

deserves space is set by how often an answer is wanted multiplied by what

producing it costs. Where none of that holds, no cache will help.

Recap

  • Requests are not spread evenly over the things that could be requested, and the gap is not small: the most popular item in a collection is often asked for thousands of times more often than the median one, which is the entire reason a small cache can carry a large share of the load.
  • What makes an answer worth keeping is the chance it is asked again multiplied by what producing it costs, so a rarely wanted answer that takes a second to build can be a better candidate than a common one that takes a microsecond.
  • Reuse is concentrated in time as well as in popularity, so most of what a cache earns comes from requests arriving within minutes of each other rather than from anything held for a long time.

This is the reading half

Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.

The contents

NextWhat a Hit Is Actually Worth →

The rest of this course

  1. 01The Same Thing, Asked Againyou are here
  2. 02What You Actually Buy With a Hitopening only
  3. 03Choosing What to Forgetopening only
  4. 04When the Copy Stops Being Trueopening only
  5. 05Choosing How Far Away to Keep Itopening only
  6. 06The Ways a Cache Turns On Youopening only
  7. 07Many Boxes Pretending to Be Oneopening only
  8. 08Telling Whether It Earns Its Placeopening only

Read alongside