ContentsThe library

Cutting Text Into Pieces

Two Bills That Move in Opposite Directions

Last timeBytes Underneath Everything

A bigger vocabulary costs parameters and makes every sequence shorter. A smaller one does the reverse. Adding the two bills gives a curve with a bottom, and that bottom is the answer.

What each entry costs

A vocabulary entry is not a word in a list. It is a row of numbers.

Going in, each entry needs a vector, so the input table is the vocabulary size

times the vector width. Coming out, the model produces a score for every entry,

so the output layer is the same shape again. A vocabulary of fifty thousand

entries on a model with vectors two thousand numbers wide is therefore two

hundred million parameters before any of the model proper exists.

That is the first cost, and the important thing about it is its shape: it is a

straight line. Doubling the vocabulary doubles it, exactly, with no diminishing

anything. On a small model this is most of the parameter budget. On a very large

one it is a few per cent, which is the main reason bigger models can afford

bigger vocabularies.

FIG 1What real models chose
vocabulary entriespieces per English wordvocabulary as per cent o
an early transformer32000.001.3512.00
the GPT-2 generation50257.001.3116.00
a mid-size open model32000.001.404.00
a recent multilingual mo128000.001.225.00
a very large multilingua256000.001.153.00
Approximate figures, and the pattern is what matters. The two marked cells are the extremes of the third column: on the smaller older model the vocabulary was a sixth of the whole thing, and on the largest it is three per cent, which is what lets it carry eight times as many entries.

What length costs

The second cost is compute, and compute is charged by the piece. Training cost

is roughly proportional to the number of pieces processed, and serving cost is

charged per piece directly. So anything that reduces pieces per word reduces

both bills in the same proportion.

A larger vocabulary does exactly that, because more merges have been learned, so

more text is covered by whole symbols. But it does so along a flattening curve.

The first few thousand merges take the rate from six pieces per word down to

below two. The next hundred thousand take it from 1.3 to 1.15. There is a floor

near one piece per word that no vocabulary reaches, because the long tail of

rare words never becomes common enough to merge.

FIG 2The two bills and their sum
0.0025.0050.0075.00100.004.033.062.091.0120.0vocabulary size, in thousands of entries
parameters: a straight linecompute: falls and flattensthe total
The total is the only curve that matters and it has a bottom, at around thirty-three thousand entries for these numbers. Note how flat it is on either side: anything between twenty and sixty thousand is within a few per cent of the best, which is why nobody agonises over the exact figure.

The two pull opposite ways

It is worth stating plainly why neither cost can pick the answer alone, because

both arguments get made in isolation and both are wrong.

Minimise parameters and the answer is 256: the byte vocabulary, no merges at

all, the smallest possible input table. It is also six times longer on every

sequence, so the compute bill rises by more than six and the saving is wiped out

many times over.

Minimise length and the answer is one entry per word in the language, plus one

per inflection, plus one per name. That is millions of entries, an input table

larger than the model, an output layer that dominates every forward pass, and

most of those entries seen so rarely during training that their vectors never

become useful.

The lesson stops here

2 more paragraphs to go

You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.

The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.

See the planThe contents

This is the reading half

Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.

The contents

The rest of this course

  1. 01Two Obvious Vocabularies, Both of Which Fail
  2. 02Count the Pairs, Join the Winner, Repeatopening only
  3. 03Working From Bytes Makes the Vocabulary Totalopening only
  4. 04Two Bills That Move in Opposite Directionsyou are here
  5. 05The Space Belongs to the Word That Follows Itopening only
  6. 06Where the Digits Get Cut Decides Whether the Sum Worksopening only
  7. 07The Same Sentence, Nine Times the Billopening only
  8. 08The One Decision That Is Made Before Anything Elseopening only

Read alongside