ContentsThe library

What It Costs to Run a Model

The Ceiling You Cannot Code Around

Last timeHow Many Bits a Number Needs

One request writing one word at a time must read the whole model from memory for every word, and that alone fixes the fastest it can possibly go.

Here is the central fact of the course, and it is worth stating before any of the

reasoning. To produce one word, a model reads all of its weights. Not some of

them, not the relevant ones. Every table in every layer is multiplied by the

numbers passing through, so every weight has to make the journey from memory to

the multipliers, and it has to make it again for the next word.

That single sentence, combined with the bandwidth figure from the first lesson,

settles how fast a conversation can possibly go.

FIG 1The fastest one request can be answered
the number of weights in the model
the bytes each weight occupies in whatever format it is stored in
the bytes the memory can deliver each second
words a second, at most, for a single request producing them one at a time
A seventy billion weight model at two bytes a weight is a hundred and forty gigabytes. Three thousand gigabytes a second delivers that in about forty-seven milliseconds, which is twenty-one words a second and not one more.

Checking that the arithmetic really is irrelevant

It is easy to accept this too quickly, so it is worth doing the other half of the

sum. Every weight contributes one multiplication and one addition, so a model of

seventy billion weights performs about a hundred and forty billion operations to

produce a word. A chip advertising three hundred trillion operations a second

gets through that in under half a millisecond.

FIG 2Both costs, for four model sizes
stepweights, in billionsgigabytes read per wordmilliseconds of movementmilliseconds of arithmeticwhat happened
17144.70.05A small model, comfortably faster than anyone reads. The arithmetic is already a rounding error.
213268.70.09Doubling the weights doubles both columns, so the ratio between them never moves.
37014046.70.47The common serving size. Twenty-one words a second, with the multipliers busy one percent of the time.
4400800266.72.7A frontier size on one machine. Under four words a second, which is why models this large are always spread over many chips.
4 steps
The last two columns differ by a factor of a hundred in every row, because both are proportional to the number of weights and the ratio between them belongs to the hardware rather than to the model.

Forty-seven milliseconds against half a millisecond. The chip's headline figure,

the one it is sold on, governs one percent of the time taken. If a vendor doubled

its arithmetic throughput tomorrow and left the memory alone, a single

conversation would run at exactly the same speed.

The lesson stops here

4 more paragraphs to go

You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.

The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.

See the planThe contents

This is the reading half

Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.

The contents

The rest of this course

  1. 01The Two Kinds of Work
  2. 02Paying for Precisionopening only
  3. 03The Ceiling You Cannot Code Aroundyou are here
  4. 04What the Model Holds On Toopening only
  5. 05The Length at Which Everything Changesopening only
  6. 06The One Trick That Actually Worksopening only
  7. 07Reading and Writing Are Different Jobsopening only
  8. 08What a Million Words Actually Costsopening only

Read alongside