Three and a Half of the Four Seconds Are Invisible, Because the Profiler Only Looks When the Program Is Running
Last timeReading the Shape
A processor profile samples a thread while it executes. A thread that is waiting executes nothing, so the wait leaves no trace and the profile reports a program with no problem.
What a processor profile misses
The mechanism from the second lesson has a consequence that is obvious once
seen and surprises almost everybody the first time.
A sampling profiler interrupts a thread while it is executing and records the
stack. A thread that is blocked, waiting for a disk read or a network reply or
a lock somebody else holds, is not executing. It is not interrupted. It
produces no samples.
So a request that spends half a second computing and three and a half seconds
waiting produces a profile of half a second. The profile is correct, complete
and useless for the question being asked, and it will typically look healthy:
time spread reasonably across a handful of functions, no obvious hot spot.
| step | phase | wall clock, seconds | processor samples | what the profile shows | what happened |
|---|---|---|---|---|---|
| 1 | parse the input | 0.2 | 200 | a visible entry | Real computation. This appears in the profile at roughly its true size relative to the other computing phases. |
| 2 | wait for the database | 2.4 | 0 | nothing at all | Sixty per cent of the request and zero samples. The thread was blocked the entire time, so the profiler never looked at it. There is no line for this anywhere in the output. |
| 3 | wait for a lock | 1.1 | 0 | nothing at all | Another twenty seven per cent, equally invisible. Worse, the holder of the lock is a different thread doing real work, so its samples make the program look busy rather than stuck. |
| 4 | format the reply | 0.3 | 300 | the largest entry | The profile reports this as the biggest cost in the request, and relative to the other sampled work it is. It is seven per cent of what the user waited for. |
The states a thread is in
The underlying model is small and worth knowing precisely, because the names
appear in every tool.
The third state deserves a note because it is misread so often. A thread that
is ready but not running is not waiting for anything in your program. It is
waiting for a processor, because the machine has more runnable work than cores.
Adding threads makes it worse, optimising code barely helps, and the symptom is
that everything is slow at once rather than one operation being slow.
The lesson stops here
4 more paragraphs to go
You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.
The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.
See the planThe contentsThis is the reading half
Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.
The contents