The Function at Number Seven Received Nine Samples, Which Means You Know Almost Nothing About It
Last timeWatching Sometimes, or Counting Everything
Every share in a sampling profile is an estimate with an uncertainty, and the uncertainty depends on how many samples that line received. One square root tells you which lines to believe.
The share is an estimate
A sampling profile looks like a table of facts. It is a table of estimates, and
each one came from a finite number of observations.
The mechanism makes this obvious once stated. The profiler took three thousand
looks at the program. Each look either found a particular function on the stack
or did not. The share reported for that function is the number of times it was
found divided by three thousand. That is a proportion estimated from a sample,
which is the same object a survey produces, and it carries the same kind of
uncertainty for the same reason.
- standard error on the estimated share
- the share the profile reports
- total samples in the run
Computing the uncertainty
The formula above is exact and slightly awkward to apply by eye. There is a
rearrangement that is easier to use and is the one worth carrying around.
For a small share, the relative error on a line is approximately one divided by
the square root of the number of samples that line received. Not the total
samples in the run: the samples on that line.
| samples received | share, per cent | plus or minus, points | relative error, per cent | |
|---|---|---|---|---|
| the parser | 1050.00 | 35.00 | 0.87 | 2.00 |
| the hash lookup | 400.00 | 13.30 | 0.62 | 5.00 |
| the retry path | 60.00 | 2.00 | 0.26 | 13.00 |
| the logger | 9.00 | 0.30 | 0.10 | 33.00 |
| the configuration read | 2.00 | 0.07 | 0.05 | 71.00 |
The practical consequence is a reading habit. Before comparing two lines of a
profile, look at how many samples each received. Two entries at 1.1 and 0.9 per
cent in a three thousand sample run received about 33 and 27 samples, with a
relative error near twenty per cent each, so they are indistinguishable and the
order they appear in is an accident of this particular run.
The lesson stops here
2 more paragraphs to go
You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.
The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.
See the planThe contentsThis is the reading half
Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.
The contents