It Is Slow Is Not a Measurement, and Profiling Before You Have One Wastes the Week
Four different complaints all arrive as the word slow, and each one needs a different measurement. Turning the complaint into a statement with a number in it is the first and largest step.
What the complaint means
Performance work begins with somebody saying a thing is slow, and that sentence
contains no measurement. Before any profiler runs, it has to be converted into
a statement that could be true or false.
Four questions do the conversion. Who waited. For which operation. How long.
And how often does it happen.
| reports per week | share of requests involv | seconds before they call | |
|---|---|---|---|
| the dashboard is slow | 3 | 100 | 4 |
| exports keep timing out | 12 | 2 | 45 |
| everything got slower af | 40 | 100 | 1 |
| the app feels laggy some | 6 | 8 | 2 |
The second row deserves attention. Two per cent of requests and twelve reports
a week is a real problem affecting real people, and it is invisible in every
aggregate measurement anyone would naturally take. The third row is the
opposite: it concerns everything, which makes it easy to measure and usually
means a change in configuration rather than in code.
The fourth question, how often, is the one most often skipped and it decides
the whole shape of the investigation. A problem that happens always can be
reproduced and profiled in an afternoon. A problem that happens two per cent of
the time requires capturing the slow cases specifically, which is a different
technique with different tools, and starting it by profiling a typical request
produces a clean profile of something that is working fine.
Latency or throughput
The second separation is between how long one operation takes and how many
operations finish per unit time. These sound like the same statement in two
forms and they are not.
- operations in progress at any moment
- operations completing per second
- how long one operation takes
The relation also shows why the two can move in opposite directions. Batching
work increases the time any individual item waits, because it sits in the batch
until the batch is full, and increases the number completing per second,
because the fixed cost is shared. A change described as making the system
faster can therefore make every user wait longer.
Which one the complaint is about is usually clear once asked. A person waiting
at a screen is a latency complaint. A queue that is not draining, a nightly job
that no longer finishes by morning, a cost per request that is too high: those
are throughput. Asking before measuring avoids the common failure of improving
the number nobody was complaining about.
Which percentile
The third separation is where in the distribution to look, and it is where most
measurement goes wrong.
Averages fail here for a structural reason. Request time distributions have a
floor and no ceiling, so they are stretched to the right, and a mean over such
a distribution is pulled upward by the tail without representing it. Reporting
the mean satisfies nobody: it is higher than what most users see, so it
overstates the typical experience, and far lower than what complaining users
see, so it understates the problem.
Two further points about tails are worth carrying. They compound. If a page
needs five service calls and each has a one in a hundred chance of being slow,
the page is slow about five per cent of the time, so a tail that looks
negligible per service is common per page. And the tail is where the
interesting causes live: queueing, garbage collection, a cache miss, a retry, a
lock held by something else. The median tells you about the common path, which
is usually the path that is working.
Choosing the measurement
With the complaint stated, the measurement chooses itself, and the test for
whether you have chosen correctly is to write down the two outcomes in advance.
The next lesson takes the first real choice of method: watching the program at
intervals, or counting every single thing it does. They have different
overheads, different resolutions and different distortions, and the choice is
determined by the question written down here.
Recap
- A complaint about speed is a statement about somebody waiting, and the first job is to find out who, for what, and how often, because those three answers choose the measurement.
- Latency and throughput are different quantities with different remedies, and a system can be fixed on one while getting worse on the other.
- Name the percentile before you measure. An average hides exactly the behaviour people complain about, and the complaint is nearly always about the tail.
This is the reading half
Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.
The contents