Two Hundred Times the Price, for a Line That Looks the Same
Last timeTwo Worlds and a Door
A request across the boundary costs hundreds of times what a function call costs, and the reasons are specific. Here is where the time goes, and how to pay it once instead of a thousand times.
What the crossing actually does
Nothing in the syntax warns you. A function call and a request to the kernel
look alike in every language, and the second one costs between one and three
hundred times the first.
The reason is that a crossing is not a jump. It is a sequence of fixed work,
performed whether the request turns out to be large or trivial.
The register state has to be saved, because the kernel is about to use the same
processor and cannot disturb yours. The stack has to change, because your stack
is in memory you control and the kernel will not run on memory a program can
write under it. The privilege state changes. On most machines since 2018 the
address mappings change too, for reasons the third section covers. The
arguments get validated, every pointer checked against what your process is
allowed to touch. And on the way back, all of it is undone.
That list is the cost, and the important thing about it is that none of the
items depend on what you asked for. Reading one byte and reading a megabyte pay
the same crossing.
| step | step | what happens | nanoseconds | running total | what happened |
|---|---|---|---|---|---|
| 1 | 1 | your code puts a number and arguments in | 2 | 2 | Ordinary instructions in the weaker state. This part really is as cheap as it looks. |
| 2 | 2 | the trap instruction, state switch, mapp | 180 | 182 | The fixed cost. On a machine without the 2018 mitigations this step is nearer sixty. |
| 3 | 3 | the kernel validates the arguments and l | 60 | 242 | Every pointer checked against the calling process. Cheap individually, unavoidable, and paid every time. |
| 4 | 4 | the work: reading from a buffer the kern | 90 | 332 | The only step that is about your request. If the data were not already in memory this would be a hundred thousand instead of ninety. |
| 5 | 5 | return, undo everything from step two | 170 | 502 | The exit is nearly as expensive as the entry, because it is the same work in reverse. |
The ladder of costs
Performance reasoning without these numbers is guessing, and most people are
out by at least two orders of magnitude somewhere on this list.
| nanoseconds | same number, so the scal | |
|---|---|---|
| one add | 1 | 1 |
| a function call | 2 | 2 |
| a miss to main memory | 20 | 20 |
| a request to the kernel | 500 | 500 |
| a minor page fault | 3000 | 3000 |
| switching to another pro | 16000 | 16000 |
| reading from a solid-sta | 100000 | 100000 |
| reading from a spinning | 10000000 | 10000000 |
Two entries deserve comment. The page fault at three thousand nanoseconds is
what it costs when your program touches memory the kernel has not yet given it
a real page for, which happens far more often than people expect and is the
subject of the memory course on this shelf. And the process switch at sixteen
thousand is the number that makes threads look cheap and processes look
expensive, which is only half true and is the subject of a later lesson.
The lesson stops here
3 more paragraphs to go
You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.
The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.
See the planThe contentsThis is the reading half
Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.
The contents