Everything You Remember Has to Fit in the Same Box
A model can remember by keeping a block of numbers of fixed size and rewriting it at every step. The size never grows, which is the whole advantage and the whole problem.
A model reading a sequence has to use what came before. There are exactly two
ways to do it, and nearly every design is one of them, the other, or a mixture.
The two ways
This course is about the right-hand branch. The block of numbers is called the
state, and everything that follows is about what it can hold.
- the state after step t: a fixed block of numbers, the same size at every step
- the state from the previous step, which is the only channel through which anything older can arrive
- the current input, the one new piece of information available at this step
- the update rule, which is the same function at every step and has no idea how far into the sequence it is
What a state update is allowed to see
The update rule is the same at every step. That is not a simplification made for
convenience: it is what lets the model handle a sequence of any length, since a
rule that depended on the position would have nothing to say about position ten
thousand if it had only been trained up to five hundred.
This is worth sitting with, because it is unlike how people describe memory to
themselves. You do not have a note that says what the third sentence was. You
have one block of numbers, and the third sentence influenced it, and the fourth
sentence influenced it again on top of that, and so on. If you want to know what
the third sentence said, there is nowhere to look. There is only the block, and
whatever of the third sentence is still detectable inside it.
Watching a state change
Here is a state of three numbers being updated over four steps, with an update
rule that partly keeps what was there and partly folds in the new input.
| step | first slot | second slot | third slot | what happened |
|---|---|---|---|---|
| 1 | 0 | 0 | 0 | The state before anything has been read. Nothing is remembered because nothing has happened. |
| 2 | 0.6 | -0.2 | 0.1 | After the first input. The block now carries some trace of it, spread across all three slots rather than stored in one. |
| 3 | 0.5 | 0.4 | -0.3 | After the second. The first input has not been erased, but every slot has moved, so what remains of it is now mixed with the new arrival. |
| 4 | 0.4 | 0.3 | 0.5 | After the third. Each update shrinks the contribution of everything older, which is the behaviour that becomes a serious problem over hundreds of steps. |
The size is the limit
Three numbers obviously cannot summarise a long document. Neither can three
thousand, for some tasks. The point is not the particular size but that it is
fixed in advance, so the amount of detail that can survive is decided before the
model ever sees how long the sequence will be.
A useful way to think about the limit is to ask what a task needs recalled. If
the question is roughly what this document was about, a summary is the right
shape of answer and a fixed block is a reasonable place to keep it. If the
question is what number appeared in the fourth line of a table nine thousand
words ago, a summary is the wrong shape, because that number was never going to
be important enough at the time to earn space that everything else was competing
for. Tasks of the second kind are exactly where carried state does badly and
where keeping every position earns its cost.
| after 10 items | after 100 | after 1000 | after 10000 | |
|---|---|---|---|---|
| carrying a fixed state | 512 | 512 | 512 | 512 |
| keeping every position | 128 | 1280 | 12800 | 128000 |
The flat line is also why the two approaches get mixed in practice. A design
that carries state for most of its layers and looks back over a short recent
window in a few of them gets the flat memory curve for the bulk of the work and
exact recall of the recent past where it matters most. Several of the systems
that work well now have that shape, and the reason is visible in these two
pictures rather than in anything subtle.
What to hold on to
A state is a fixed block of numbers that stands in for everything that came
before. The update at each step sees only the previous state and the current
input, which means the far past reaches the present only by having survived
every update in between. The size of the block never changes, so the work per
item never changes and the fidelity of the summary must. Whether that trade is
acceptable is a question about the task, and the rest of this course is about
making the surviving half as useful as possible.
Recap
- State is a fixed-size block of numbers that stands in for the entire past, rewritten at each step from its previous value and the current input, and nothing else from the past is available.
- Because the block never grows, the cost of handling one more item is the same whether it is the second item or the millionth, which is the property that makes this approach cheap to run.
- The same fixed size is the limit: a summary of a thousand items and a summary of ten items occupy the same space, so something has to be discarded, and which things get discarded is the entire design problem.
This is the reading half
Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.
The contents