ContentsThe library

Reading a Profile Honestly

It Is Slow Is Not a Measurement, and Profiling Before You Have One Wastes the Week

Four different complaints all arrive as the word slow, and each one needs a different measurement. Turning the complaint into a statement with a number in it is the first and largest step.

What the complaint means

Performance work begins with somebody saying a thing is slow, and that sentence

contains no measurement. Before any profiler runs, it has to be converted into

a statement that could be true or false.

Four questions do the conversion. Who waited. For which operation. How long.

And how often does it happen.

FIG 1Four complaints that all arrived as the word slow
reports per weekshare of requests involvseconds before they call
the dashboard is slow31004
exports keep timing out12245
everything got slower af401001
the app feels laggy some682
The same word in all four cases and four completely different investigations. The marked cells are the interesting ones: two of these complaints concern a small minority of requests, which means an average over all requests will show nothing at all and the investigation will conclude there is no problem.

The second row deserves attention. Two per cent of requests and twelve reports

a week is a real problem affecting real people, and it is invisible in every

aggregate measurement anyone would naturally take. The third row is the

opposite: it concerns everything, which makes it easy to measure and usually

means a change in configuration rather than in code.

The fourth question, how often, is the one most often skipped and it decides

the whole shape of the investigation. A problem that happens always can be

reproduced and profiled in an afternoon. A problem that happens two per cent of

the time requires capturing the slow cases specifically, which is a different

technique with different tools, and starting it by profiling a typical request

produces a clean profile of something that is working fine.

Latency or throughput

The second separation is between how long one operation takes and how many

operations finish per unit time. These sound like the same statement in two

forms and they are not.

FIG 2The relation that connects them
operations in progress at any moment
operations completing per second
how long one operation takes
A system completing 200 requests a second with each taking half a second has 100 in flight. The relation holds regardless of what the system does internally, which makes it the most useful single equation in performance work: it lets any one of the three be computed from the other two, and the one people never measure is usually the one that explains the problem.

The relation also shows why the two can move in opposite directions. Batching

work increases the time any individual item waits, because it sits in the batch

until the batch is full, and increases the number completing per second,

because the fixed cost is shared. A change described as making the system

faster can therefore make every user wait longer.

Which one the complaint is about is usually clear once asked. A person waiting

at a screen is a latency complaint. A queue that is not draining, a nightly job

that no longer finishes by morning, a cost per request that is too high: those

are throughput. Asking before measuring avoids the common failure of improving

the number nobody was complaining about.

Which percentile

The third separation is where in the distribution to look, and it is where most

measurement goes wrong.

FIG 3One day of request times
nine out of ten requests, nobody complainsthe tenth, and every complaint you have received050010002000180median240mean420p901500p99milliseconds to complete a request
The mean sits at 240 and describes nobody: it is above most requests and far below the ones people notice. The p99 at 1500 is eight times the median, which is ordinary rather than pathological. All of the complaints come from the right-hand span and none of them would move the mean enough to see.

Averages fail here for a structural reason. Request time distributions have a

floor and no ceiling, so they are stretched to the right, and a mean over such

a distribution is pulled upward by the tail without representing it. Reporting

the mean satisfies nobody: it is higher than what most users see, so it

overstates the typical experience, and far lower than what complaining users

see, so it understates the problem.

Two further points about tails are worth carrying. They compound. If a page

needs five service calls and each has a one in a hundred chance of being slow,

the page is slow about five per cent of the time, so a tail that looks

negligible per service is common per page. And the tail is where the

interesting causes live: queueing, garbage collection, a cache miss, a retry, a

lock held by something else. The median tells you about the common path, which

is usually the path that is working.

Choosing the measurement

With the complaint stated, the measurement chooses itself, and the test for

whether you have chosen correctly is to write down the two outcomes in advance.

FIG 4From a complaint to a measurement
Four decisions, none of which require a profiler, and all of which change what the profiler should be pointed at. The step at the bottom before measurement is the one professionals add and amateurs skip.
FIG 5What the problem turned out to be, across investigated complaints
Only a fifth of these were the thing a processor profile is good at finding. A third were the program waiting, where a processor profile shows an idle program and reports nothing wrong. This distribution is the reason the course spends its first half on what to measure and only then on how to read it.

The next lesson takes the first real choice of method: watching the program at

intervals, or counting every single thing it does. They have different

overheads, different resolutions and different distortions, and the choice is

determined by the question written down here.

Recap

  • A complaint about speed is a statement about somebody waiting, and the first job is to find out who, for what, and how often, because those three answers choose the measurement.
  • Latency and throughput are different quantities with different remedies, and a system can be fixed on one while getting worse on the other.
  • Name the percentile before you measure. An average hides exactly the behaviour people complain about, and the complaint is nearly always about the tail.

This is the reading half

Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.

The contents

NextWatching Sometimes, or Counting Everything →

The rest of this course

  1. 01It Is Slow Is Not a Measurement, and Profiling Before You Have One Wastes the Weekyou are here
  2. 02One Profiler Interrupts a Thousand Times a Second and Guesses, the Other Watches Every Call and Changes the Answeropening only
  3. 03The Function at Number Seven Received Nine Samples, Which Means You Know Almost Nothing About Itopening only
  4. 04The Function Holding Ninety Per Cent of the Time Does Nothing, and the One Doing the Work Holds Four Per Centopening only
  5. 05The Same Function Is Cheap in Two Places and Ruinous in the Third, and the Flat List Shows One Numberopening only
  6. 06Three and a Half of the Four Seconds Are Invisible, Because the Profiler Only Looks When the Program Is Runningopening only
  7. 07The Sample Landed Three Instructions After the One That Was Slow, and the Function It Blames Does Not Exist Any Moreopening only
  8. 08The New Version Is Eight Per Cent Faster, and So Is the Old One If You Run It Enough Timesopening only

Read alongside