Working From Bytes Makes the Vocabulary Total
Last timeMerging the Commonest Pair
There are too many characters in the world to list, so modern tokenizers start from the 256 bytes instead. That makes nothing unrepresentable, and it makes some scripts very expensive.
Characters are not a workable alphabet
The previous two lessons said that every single character stays in the
vocabulary as a fallback, so that nothing is ever unrepresentable. That works
nicely for a corpus of English. It does not survive contact with the world.
Unicode defines close to 150,000 characters, and more arrive with every
revision. Reserving one vocabulary entry for each would consume five times the
entire budget of a typical model before a single useful piece had been learned.
Reserving entries only for the characters that appear in the corpus brings back
exactly the hole the fallback was meant to close: the first document containing
a script the corpus lacked is unrepresentable, and that document certainly
exists.
So characters cannot be the atoms. Something smaller is needed, and there is
only one candidate.
Bytes make it total
Every file, every message, every piece of text that has ever been transmitted
is a sequence of bytes, and there are exactly 256 of them. Put all 256 in the
vocabulary and the guarantee is absolute: every possible input, in every script,
including malformed text and binary data that is not text at all, has a
representation. The worst case is that a string is spelled out one byte at a
time, which is long but never impossible.
| step | the text | characters | bytes | pieces under an English vocabulary | what happened |
|---|---|---|---|---|---|
| 1 | the cat | 7 | 7 | 2 | English. One byte per character, and the words are common enough that the vocabulary has them whole. |
| 2 | le chat | 7 | 7 | 3 | French in plain Latin letters. Same bytes, slightly more pieces, because the corpus contained less French. |
| 3 | the Russian for cat | 3 | 6 | 4 | Three Cyrillic characters, two bytes each, and the vocabulary has learned few Cyrillic merges, so six bytes become four pieces. |
| 4 | the Hindi for cat | 6 | 18 | 12 | Six Devanagari characters at three bytes each, almost no learned merges, so the result sits close to the byte floor. Six times the English piece count for the same word. |
This is the step that makes the earlier guarantee true rather than
approximately true, and it is why current tokenizers are described as byte
level. The merging algorithm is unchanged. It simply starts from 256 symbols
rather than from whatever characters happened to appear.
The lesson stops here
4 more paragraphs to go
You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.
The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.
See the planThe contentsThis is the reading half
Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.
The contents