The library

How a program runs

Follow one memory access from a variable in your code to the chip that answers it, well enough to predict which accesses are fast, to explain a hundredfold difference between two loops that do the same arithmetic, and to lay data out on purpose

Memory, All the Way Down

An address in your program is not an address in the hardware, and the distance between them is where most performance lives. This course follows one access down through the address translation, the caches and the physical memory.

8 lessons, written and corrected before you arrived. Reading them here needs no account. The first reads the whole way through; the others open and then stop, because a page nobody owns cannot tell who is reading it. Starting the course gives you your own copy, where every idea has problems standing under it and you can ask about any sentence.

Start reading

  1. 01Two Programs Can Hold the Same Address and Never CollidePrint a pointer in two programs and you can get the same number twice. Neither is lying and neither is the real address. Here is what sits in between.
  2. 02A Few Hundred Entries Decide Whether Your Program Falls Off a Cliffopening onlyTranslation would cost four memory reads per access if it were not cached. The cache is tiny, it covers a fixed amount of memory, and programs that outgrow it slow down suddenly.
  3. 03Most Faults Cost a Microsecond and One Costs Ten Thousandopening onlyA missing block interrupts your program and hands control to the system. Four things can happen next, three of them cheap, and conflating them is why fault counts get ignored.
  4. 04If a Cache Hit Took One Second, Main Memory Would Take Four Minutesopening onlyThe gaps in the memory hierarchy are two orders of magnitude and nobody can feel a nanosecond. Rescale the whole ladder to seconds and the design rules become obvious.
  5. 05You Asked for Four Bytes and Sixty-Four Arrivedopening onlyMemory is not moved a byte at a time. It moves in blocks of sixty-four, and whether you use the other sixty is the difference between fast code and slow code.
  6. 06Predict the Speedup on Paper Before You Change a Lineopening onlySplitting one array of records into several arrays of fields is the single highest-value layout change, and the gain is calculable from two numbers before you write any code.
  7. 07One Allocator Adds a Number, the Other Goes Lookingopening onlyThe two places a program puts things differ by a factor of a hundred in cost, by everything in failure mode, and the choice between them is about lifetime rather than speed.
  8. 08Two Threads, No Shared Variables, and One of Them Is Ten Times Sloweropening onlyCaches must agree, so a write by one core takes a block away from another. Two threads that share nothing can still fight, purely over where their data happens to sit.