ContentsThe library

Models That Carry a Memory

Everything You Remember Has to Fit in the Same Box

A model can remember by keeping a block of numbers of fixed size and rewriting it at every step. The size never grows, which is the whole advantage and the whole problem.

A model reading a sequence has to use what came before. There are exactly two

ways to do it, and nearly every design is one of them, the other, or a mixture.

FIG 1Looking back against carrying forward
Neither branch is a compromise of the other. They sit at opposite ends of a trade between what you keep and what you pay, and the correct choice depends entirely on whether the task needs exact recall of specific earlier items.

The two ways

This course is about the right-hand branch. The block of numbers is called the

state, and everything that follows is about what it can hold.

FIG 2The whole of the idea, in one line
the state after step t: a fixed block of numbers, the same size at every step
the state from the previous step, which is the only channel through which anything older can arrive
the current input, the one new piece of information available at this step
the update rule, which is the same function at every step and has no idea how far into the sequence it is
Note what is absent. There is no term for the input ten steps ago and no term for how long the sequence has been running. Everything the model will ever know about step one must have survived every update since, which is the constraint that the rest of this course keeps running into.

What a state update is allowed to see

The update rule is the same at every step. That is not a simplification made for

convenience: it is what lets the model handle a sequence of any length, since a

rule that depended on the position would have nothing to say about position ten

thousand if it had only been trained up to five hundred.

This is worth sitting with, because it is unlike how people describe memory to

themselves. You do not have a note that says what the third sentence was. You

have one block of numbers, and the third sentence influenced it, and the fourth

sentence influenced it again on top of that, and so on. If you want to know what

the third sentence said, there is nowhere to look. There is only the block, and

whatever of the third sentence is still detectable inside it.

Watching a state change

Here is a state of three numbers being updated over four steps, with an update

rule that partly keeps what was there and partly folds in the new input.

FIG 3Three numbers, four steps
stepfirst slotsecond slotthird slotwhat happened
1000The state before anything has been read. Nothing is remembered because nothing has happened.
20.6-0.20.1After the first input. The block now carries some trace of it, spread across all three slots rather than stored in one.
30.50.4-0.3After the second. The first input has not been erased, but every slot has moved, so what remains of it is now mixed with the new arrival.
40.40.30.5After the third. Each update shrinks the contribution of everything older, which is the behaviour that becomes a serious problem over hundreds of steps.
4 steps
No slot corresponds to an input. The first item is not in position one; it is smeared across all three and fading. That is why you cannot look at a trained state and read off what it remembers, and it is also why the capacity question is about the whole block rather than about slots.

The size is the limit

Three numbers obviously cannot summarise a long document. Neither can three

thousand, for some tasks. The point is not the particular size but that it is

fixed in advance, so the amount of detail that can survive is decided before the

model ever sees how long the sequence will be.

A useful way to think about the limit is to ask what a task needs recalled. If

the question is roughly what this document was about, a summary is the right

shape of answer and a fixed block is a reasonable place to keep it. If the

question is what number appeared in the fourth line of a table nine thousand

words ago, a summary is the wrong shape, because that number was never going to

be important enough at the time to earn space that everything else was competing

for. Tasks of the second kind are exactly where carried state does badly and

where keeping every position earns its cost.

FIG 4Numbers that must be kept after reading a sequence of a given length
after 10 itemsafter 100after 1000after 10000
carrying a fixed state512512512512
keeping every position128128012800128000
The top row is flat by construction. The bottom row is the one that forces architectural decisions at serving time, because it is per request and it never stops growing while a response is being generated. The marked cell is where keeping everything stops being free.
FIG 5How memory grows as a sequence gets longer
0.001500.003000.004500.006000.001.0100.8200.5300.3400.0items read so far
a fixed statea record of every position
The flat line is the promise of a carried state: a model serving a conversation of a hundred thousand tokens uses the same memory per request as one serving ten. The rising line crosses it early and then keeps going, and for long sequences the gap is the whole argument.

The flat line is also why the two approaches get mixed in practice. A design

that carries state for most of its layers and looks back over a short recent

window in a few of them gets the flat memory curve for the bulk of the work and

exact recall of the recent past where it matters most. Several of the systems

that work well now have that shape, and the reason is visible in these two

pictures rather than in anything subtle.

What to hold on to

A state is a fixed block of numbers that stands in for everything that came

before. The update at each step sees only the previous state and the current

input, which means the far past reaches the present only by having survived

every update in between. The size of the block never changes, so the work per

item never changes and the fidelity of the summary must. Whether that trade is

acceptable is a question about the task, and the rest of this course is about

making the surviving half as useful as possible.

Recap

  • State is a fixed-size block of numbers that stands in for the entire past, rewritten at each step from its previous value and the current input, and nothing else from the past is available.
  • Because the block never grows, the cost of handling one more item is the same whether it is the second item or the millionth, which is the property that makes this approach cheap to run.
  • The same fixed size is the limit: a summary of a thousand items and a summary of ten items occupy the same space, so something has to be discarded, and which things get discarded is the entire design problem.

This is the reading half

Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.

The contents

NextOne Step, Applied Over and Over →

The rest of this course

  1. 01Everything You Remember Has to Fit in the Same Boxyou are here
  2. 02The Same Small Rule, Run a Thousand Timesopening only
  3. 03Almost Right, Four Hundred Times in a Rowopening only
  4. 04Add Instead of Multiply, and Decide How Muchopening only
  5. 05A Thousand Waits That Did Not Have to Happenopening only
  6. 06Slots That Forget at Rates You Choseopening only
  7. 07A Rule That Reads What Arrived Before Decidingopening only
  8. 08The Bill That Grows and the Bill That Does Notopening only

Read alongside