ContentsThe library

What It Costs to Run a Model

The One Trick That Actually Works

Last timeWhen the Context Gets Long

Reading the weights once and using them for many requests at the same time is the single change that turns a movement problem into an arithmetic one.

The previous lessons have been mostly bad news: the weights must be read for

every word, the kept conversations must be read too, and arithmetic sits idle

through all of it. This lesson is the good news, and there is really only one

piece of it.

The weights do not care how many requests there are

A pass through a weight matrix takes whatever vectors you hand it and multiplies

each one against the same numbers. Hand it one vector and it reads the matrix and

does one vector's worth of arithmetic. Hand it sixty-four and it reads exactly

the same matrix and does sixty-four vectors' worth. The reading is identical.

FIG 1The operations performed per byte read
the number of weights in the matrix, each used once per vector passed through
how many requests are grouped into the pass
the arithmetic, which grows with the group
the bytes read, at two bytes a weight, which does not grow with the group
operations per byte, which comes out as the group size exactly
This is the single most useful fact in the course. The balance point of a machine, around three hundred operations per byte, is therefore also a target group size. Reach it and the arithmetic units are finally busy.

So the idle arithmetic from the third lesson has an answer, and the answer is to

stop sending one request at a time. Notice that this is not a clever algorithm or

a better format. It is the same work, rearranged so that an expensive reading is

amortised.

It is also the reason the serving of a model looks so different from the running

of one. A single reader talking to a model alone is served by a machine doing a

few percent of the arithmetic it is capable of, and nothing can be done about

that, because there is no one to share the reading with. The same machine serving

a hundred readers at once is doing useful work with nearly all of its arithmetic.

The hardware has not changed and the model has not changed, so almost the whole

difference in cost per word between a quiet system and a busy one comes from this

one effect.

The lesson stops here

4 more paragraphs to go

You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.

The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.

See the planThe contents

This is the reading half

Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.

The contents

The rest of this course

  1. 01The Two Kinds of Work
  2. 02Paying for Precisionopening only
  3. 03The Ceiling You Cannot Code Aroundopening only
  4. 04What the Model Holds On Toopening only
  5. 05The Length at Which Everything Changesopening only
  6. 06The One Trick That Actually Worksyou are here
  7. 07Reading and Writing Are Different Jobsopening only
  8. 08What a Million Words Actually Costsopening only

Read alongside