What the Model Holds On To
Last timeReading Every Weight for Every Word
A model consulting the words already seen keeps a summary of each one, and that store grows with every word until it rivals the model itself in size.
Up to now the model has been the only thing in memory, and it has been a fixed
size. That is why the ceiling in the last lesson was a clean division. This
lesson introduces the second occupant of the same memory, and unlike the model
it grows.
Why anything is kept at all
When a model produces a word, the layer that looks back over the conversation
compares the new position against a pair of summaries belonging to every earlier
position. One of the pair decides how much attention the new position pays to
the old one; the other is what gets collected if it pays any.
The important property is that those summaries depend only on the position they
belong to and the words before it. They were computed when that position was
first processed, and nothing that happens later changes them. So there are two
choices: keep them, or compute them again every time.
Computing them again is not a mild inefficiency. Producing the five hundredth
word would mean recomputing the summaries for four hundred and ninety-nine
positions, and the five hundred and first would redo all five hundred. The total
work would grow with the square of the length rather than in proportion to it,
which for a conversation of any size is simply unaffordable. So they are kept.
Every production system keeps them, and the interesting question is what that
costs.
- the two summaries kept per position, one for deciding attention and one for carrying content
- the number of layers, each of which keeps its own pair
- the width of the model, which is how many numbers each summary holds
- the fraction left after sharing summaries between heads, which is one if every head keeps its own
- the bytes each number occupies
- how many positions the conversation has reached
- the bytes held for that one conversation, which grows as it continues
Exactly what it costs
Eleven gigabytes is the number worth sitting with. The weights of that model, in
the same format, are a hundred and forty gigabytes. So thirteen conversations of
four thousand words hold as much memory as the entire model, and a machine that
comfortably fits the weights can be defeated by a dozen users.
The lesson stops here
3 more paragraphs to go
You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.
The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.
See the planThe contentsThis is the reading half
Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.
The contents