ContentsThe library

What It Costs to Run a Model

What the Model Holds On To

Last timeReading Every Weight for Every Word

A model consulting the words already seen keeps a summary of each one, and that store grows with every word until it rivals the model itself in size.

Up to now the model has been the only thing in memory, and it has been a fixed

size. That is why the ceiling in the last lesson was a clean division. This

lesson introduces the second occupant of the same memory, and unlike the model

it grows.

Why anything is kept at all

When a model produces a word, the layer that looks back over the conversation

compares the new position against a pair of summaries belonging to every earlier

position. One of the pair decides how much attention the new position pays to

the old one; the other is what gets collected if it pays any.

The important property is that those summaries depend only on the position they

belong to and the words before it. They were computed when that position was

first processed, and nothing that happens later changes them. So there are two

choices: keep them, or compute them again every time.

Computing them again is not a mild inefficiency. Producing the five hundredth

word would mean recomputing the summaries for four hundred and ninety-nine

positions, and the five hundred and first would redo all five hundred. The total

work would grow with the square of the length rather than in proportion to it,

which for a conversation of any size is simply unaffordable. So they are kept.

Every production system keeps them, and the interesting question is what that

costs.

FIG 1The store for one conversation
the two summaries kept per position, one for deciding attention and one for carrying content
the number of layers, each of which keeps its own pair
the width of the model, which is how many numbers each summary holds
the fraction left after sharing summaries between heads, which is one if every head keeps its own
the bytes each number occupies
how many positions the conversation has reached
the bytes held for that one conversation, which grows as it continues
Eighty layers, a width of eight thousand, two bytes a number and no sharing comes to about two and a half megabytes per position. A four thousand word conversation is then around eleven gigabytes, for one user.

Exactly what it costs

Eleven gigabytes is the number worth sitting with. The weights of that model, in

the same format, are a hundred and forty gigabytes. So thirteen conversations of

four thousand words hold as much memory as the entire model, and a machine that

comfortably fits the weights can be defeated by a dozen users.

The lesson stops here

3 more paragraphs to go

You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.

The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.

See the planThe contents

This is the reading half

Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.

The contents

The rest of this course

  1. 01The Two Kinds of Work
  2. 02Paying for Precisionopening only
  3. 03The Ceiling You Cannot Code Aroundopening only
  4. 04What the Model Holds On Toyou are here
  5. 05The Length at Which Everything Changesopening only
  6. 06The One Trick That Actually Worksopening only
  7. 07Reading and Writing Are Different Jobsopening only
  8. 08What a Million Words Actually Costsopening only

Read alongside