The library

How models work

Be able to explain how a model can process a sequence by carrying a fixed summary forward instead of looking back at everything, why that idea failed for twenty years, what fixed it, and when it beats looking back

Models That Carry a Memory

A model can remember by keeping a fixed-size summary and updating it one step at a time. That idea is older than attention, nearly died, and came back. This course is about why.

8 lessons, written and corrected before you arrived. Reading them here needs no account. The first reads the whole way through; the others open and then stop, because a page nobody owns cannot tell who is reading it. Starting the course gives you your own copy, where every idea has problems standing under it and you can ask about any sentence.

Start reading

  1. 01Everything You Remember Has to Fit in the Same BoxA model can remember by keeping a block of numbers of fixed size and rewriting it at every step. The size never grows, which is the whole advantage and the whole problem.
  2. 02The Same Small Rule, Run a Thousand Timesopening onlyA recurrent model is one modest function applied repeatedly. Unrolling it shows both where its generality comes from and why a long sequence turns it into an extremely deep network.
  3. 03Almost Right, Four Hundred Times in a Rowopening onlyRepeating one update rule means multiplying by roughly the same thing over and over, and the result either fades to nothing or grows past what can be represented. There is almost no middle.
  4. 04Add Instead of Multiply, and Decide How Muchopening onlyA gated update gives information a route through the chain that is addition rather than repeated transformation, and lets the model learn, per slot and per step, how much to keep.
  5. 05A Thousand Waits That Did Not Have to Happenopening onlyA recurrence insists on walking a known sequence in order, which wastes almost all available hardware. Removing the nonlinearity from inside the loop makes the same computation reorderable.
  6. 06Slots That Forget at Rates You Choseopening onlyOnce the update is linear, the carried state can be given one independent decay rate per slot, and spreading those rates across many timescales is what makes the memory work.
  7. 07A Rule That Reads What Arrived Before Decidingopening onlyA fixed decay rate treats every token the same, so the state fills with filler. Letting the current input set the keep and write amounts turns the state into something selective.
  8. 08The Bill That Grows and the Bill That Does Notopening onlyA look-back model stores something for every token it has seen, so serving cost climbs with the conversation. A carried state is one fixed block, and that difference decides what can be deployed.