Three Places You Could Cut
Last timeNothing Reads Letters
Cutting text into characters, into words, or into something in between are three different trades between the length of the list and the length of the sequence.
The cutting step needs a list of known pieces and a rule for matching text
against it. Everything about that step follows from how big the pieces are, and
there are three natural answers.
Cut at every character. Cut at every space. Or cut somewhere in between, at
boundaries chosen by looking at data rather than by any rule a person would
write. The third answer is what everything uses now, and the reason is clearest
after seeing what the other two cost.
Characters
Take the pieces to be single characters. The list is tiny, a few hundred entries
covering every letter, digit and mark anyone will type. Nothing is ever unknown,
because every string is made of characters by definition. Misspellings, names,
code and invented words all go through without special handling, and the round
trip is trivially exact.
Then look at the sequence length. An average English word is around five
characters including the space, so a hundred-word passage becomes five hundred
pieces instead of a hundred. Everything downstream scales with that number, and
the attention mechanism inside a model scales worse than linearly, so a factor of
five in length is considerably more than a factor of five in cost.
There is a second cost that matters more than it looks. The model now has to
spend capacity learning that certain character sequences form words at all, which
is effort not spent on anything else.
- the sequence length handed to the model
- the number of words in the text
- the average number of pieces per word, which the scheme decides
Words
Now take the pieces to be words, split on spaces. Sequences are as short as they
can possibly be, which is the whole attraction.
The list is the problem. A vocabulary that covers ordinary English text runs into
the hundreds of thousands of entries, each needing its own row in the table, and
the table is one of the largest objects in a small model. Worse, that enormous
list is still not enough. Names are effectively unlimited. Typos are unlimited.
New words appear constantly, and any word not on the list has to fall into a
single catch-all entry.
The lesson stops here
6 more paragraphs to go
You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.
The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.
See the planThe contentsThis is the reading half
Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.
The contents