The Number Somebody Had To Choose
Last timeLetting the Data Choose the Pieces
A longer list of pieces shortens every sequence but widens the table and the output layer, and the gain fades while the cost keeps rising.
The last lesson took the target size as given. Somebody chose it. Fifty thousand,
or a hundred thousand, or thirty two thousand. The number is set before any
training starts, it cannot be changed afterwards without building a new model,
and it is a genuine trade rather than a matter of taste.
Here is the shape of it. A longer list means more of the text is covered by whole
pieces, so any given passage becomes a shorter sequence. Everything downstream
costs something per position, so shorter sequences are cheaper everywhere at
once: less memory, less arithmetic, more room before the context limit. That is
the case for going long.
Against it: the list size is literally a dimension of two of the largest tables
in the model. The lookup table has one row per entry. The layer at the end that
scores every possible next piece has one column per entry. Both are the list size
multiplied by the model width, and they grow in a perfectly straight line.
- parameters in the two vocabulary-sized tables, the lookup table and the scoring layer
- the number of entries in the list
- the model width, meaning how many numbers represent one piece
The benefit fades
The first thousand entries are extremely productive. They absorb the most common
letter pairs and the most common short words, which appear everywhere, so they
remove positions from almost every sequence. The next thousand are less
productive. By the time the list is fifty thousand long, a new entry covers
material that appears perhaps once in a hundred thousand words, so it removes
almost nothing.
There is also a hard floor. Sequence length cannot fall below one piece per word,
because words are separated and the pieces cannot span the separation. So the
benefit side approaches a limit while the cost side keeps climbing.
The ceiling nobody expects
Memory is not usually what stops the list from growing. Learning is. Each row of
the lookup table is only adjusted when its entry actually appears in the training
text. A common entry is seen millions of times and ends up well positioned. An
entry down at the bottom of the frequency list might appear two hundred times in
the entire run.
Two hundred updates is not enough to learn a useful row. The result is an entry
that technically exists but whose vector is close to where it started, which can
be worse than having no entry at all, because material that would otherwise have
been cut into well learned fragments is instead represented by one badly learned
piece. This is the real reason lists stop at a hundred thousand or so rather than
a million.
The lesson stops here
3 more paragraphs to go
You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.
The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.
See the planThe contentsThis is the reading half
Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.
The contents