Two Bills That Move in Opposite Directions
Last timeBytes Underneath Everything
A bigger vocabulary costs parameters and makes every sequence shorter. A smaller one does the reverse. Adding the two bills gives a curve with a bottom, and that bottom is the answer.
What each entry costs
A vocabulary entry is not a word in a list. It is a row of numbers.
Going in, each entry needs a vector, so the input table is the vocabulary size
times the vector width. Coming out, the model produces a score for every entry,
so the output layer is the same shape again. A vocabulary of fifty thousand
entries on a model with vectors two thousand numbers wide is therefore two
hundred million parameters before any of the model proper exists.
That is the first cost, and the important thing about it is its shape: it is a
straight line. Doubling the vocabulary doubles it, exactly, with no diminishing
anything. On a small model this is most of the parameter budget. On a very large
one it is a few per cent, which is the main reason bigger models can afford
bigger vocabularies.
| vocabulary entries | pieces per English word | vocabulary as per cent o | |
|---|---|---|---|
| an early transformer | 32000.00 | 1.35 | 12.00 |
| the GPT-2 generation | 50257.00 | 1.31 | 16.00 |
| a mid-size open model | 32000.00 | 1.40 | 4.00 |
| a recent multilingual mo | 128000.00 | 1.22 | 5.00 |
| a very large multilingua | 256000.00 | 1.15 | 3.00 |
What length costs
The second cost is compute, and compute is charged by the piece. Training cost
is roughly proportional to the number of pieces processed, and serving cost is
charged per piece directly. So anything that reduces pieces per word reduces
both bills in the same proportion.
A larger vocabulary does exactly that, because more merges have been learned, so
more text is covered by whole symbols. But it does so along a flattening curve.
The first few thousand merges take the rate from six pieces per word down to
below two. The next hundred thousand take it from 1.3 to 1.15. There is a floor
near one piece per word that no vocabulary reaches, because the long tail of
rare words never becomes common enough to merge.
The two pull opposite ways
It is worth stating plainly why neither cost can pick the answer alone, because
both arguments get made in isolation and both are wrong.
Minimise parameters and the answer is 256: the byte vocabulary, no merges at
all, the smallest possible input table. It is also six times longer on every
sequence, so the compute bill rises by more than six and the saving is wiped out
many times over.
Minimise length and the answer is one entry per word in the language, plus one
per inflection, plus one per name. That is millions of entries, an input table
larger than the model, an output layer that dominates every forward pass, and
most of those entries seen so rarely during training that their vectors never
become useful.
The lesson stops here
2 more paragraphs to go
You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.
The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.
See the planThe contentsThis is the reading half
Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.
The contents