ContentsThe library

How Text Becomes Numbers

The Number Somebody Had To Choose

Last timeLetting the Data Choose the Pieces

A longer list of pieces shortens every sequence but widens the table and the output layer, and the gain fades while the cost keeps rising.

The last lesson took the target size as given. Somebody chose it. Fifty thousand,

or a hundred thousand, or thirty two thousand. The number is set before any

training starts, it cannot be changed afterwards without building a new model,

and it is a genuine trade rather than a matter of taste.

Here is the shape of it. A longer list means more of the text is covered by whole

pieces, so any given passage becomes a shorter sequence. Everything downstream

costs something per position, so shorter sequences are cheaper everywhere at

once: less memory, less arithmetic, more room before the context limit. That is

the case for going long.

Against it: the list size is literally a dimension of two of the largest tables

in the model. The lookup table has one row per entry. The layer at the end that

scores every possible next piece has one column per entry. Both are the list size

multiplied by the model width, and they grow in a perfectly straight line.

FIG 1What the list costs in parameters
parameters in the two vocabulary-sized tables, the lookup table and the scoring layer
the number of entries in the list
the model width, meaning how many numbers represent one piece
The factor of two is there because the same size appears at both ends of the model. Some models share one table between the two ends, which removes the factor and is a common economy in smaller models. At a width of 768 and fifty thousand entries that is about 77 million parameters doing nothing but converting between text and numbers.
FIG 2The cost side has no curve in it
0.00125.00250.00375.00500.001.064.8128.5192.3256.0entries in the list, in thousands
both tables, at a width of 768one shared table, at a width of 768
Straight lines, because each entry costs exactly one row and one column regardless of how useful it is. There is no point at which further entries become cheaper. Whatever benefit a longer list brings has to keep pace with this, and the next figure shows that it does not.

The benefit fades

The first thousand entries are extremely productive. They absorb the most common

letter pairs and the most common short words, which appear everywhere, so they

remove positions from almost every sequence. The next thousand are less

productive. By the time the list is fifty thousand long, a new entry covers

material that appears perhaps once in a hundred thousand words, so it removes

almost nothing.

There is also a hard floor. Sequence length cannot fall below one piece per word,

because words are separated and the pieces cannot span the separation. So the

benefit side approaches a limit while the cost side keeps climbing.

FIG 3Pieces per word, as the list lengthens
020406020406080100entries in the list, in thousands
pieces per word
Turn the knob up to represent unusual text: a specialist field, an underrepresented language, heavy code or numbers. The curve does not change shape, it simply sits higher everywhere, which means a longer list helps unusual text more than ordinary text but never brings it down to the same floor. Ordinary text is already close to the floor by thirty thousand entries, which is why that region is where most choices land.
FIG 4Where the parameters go in a small model
In a small model with a long list, a third of the parameter budget can go on converting between text and numbers rather than on anything that could be called reasoning. The same list in a very large model is a rounding error, which is why small models tend towards shorter lists and large ones can afford long ones. The trade is not the same at every scale.

The ceiling nobody expects

Memory is not usually what stops the list from growing. Learning is. Each row of

the lookup table is only adjusted when its entry actually appears in the training

text. A common entry is seen millions of times and ends up well positioned. An

entry down at the bottom of the frequency list might appear two hundred times in

the entire run.

Two hundred updates is not enough to learn a useful row. The result is an entry

that technically exists but whose vector is close to where it started, which can

be worse than having no entry at all, because material that would otherwise have

been cut into well learned fragments is instead represented by one badly learned

piece. This is the real reason lists stop at a hundred thousand or so rather than

a million.

The lesson stops here

3 more paragraphs to go

You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.

The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.

See the planThe contents

This is the reading half

Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.

The contents

The rest of this course

  1. 01The Layer Nobody Looks At
  2. 02Three Places You Could Cutopening only
  3. 03Nobody Wrote This Listopening only
  4. 04The Number Somebody Had To Chooseyou are here
  5. 05One Row, Looked Upopening only
  6. 06What Ended Up In The Tableopening only
  7. 07The Row Is Only The Starting Pointopening only
  8. 08Complaints That Are Really About This Layeropening only

Read alongside