ContentsThe library

Cutting Text Into Pieces

Working From Bytes Makes the Vocabulary Total

Last timeMerging the Commonest Pair

There are too many characters in the world to list, so modern tokenizers start from the 256 bytes instead. That makes nothing unrepresentable, and it makes some scripts very expensive.

Characters are not a workable alphabet

The previous two lessons said that every single character stays in the

vocabulary as a fallback, so that nothing is ever unrepresentable. That works

nicely for a corpus of English. It does not survive contact with the world.

Unicode defines close to 150,000 characters, and more arrive with every

revision. Reserving one vocabulary entry for each would consume five times the

entire budget of a typical model before a single useful piece had been learned.

Reserving entries only for the characters that appear in the corpus brings back

exactly the hole the fallback was meant to close: the first document containing

a script the corpus lacked is unrepresentable, and that document certainly

exists.

So characters cannot be the atoms. Something smaller is needed, and there is

only one candidate.

Bytes make it total

Every file, every message, every piece of text that has ever been transmitted

is a sequence of bytes, and there are exactly 256 of them. Put all 256 in the

vocabulary and the guarantee is absolute: every possible input, in every script,

including malformed text and binary data that is not text at all, has a

representation. The worst case is that a string is spelled out one byte at a

time, which is long but never impossible.

FIG 1The same meaning, four ways
stepthe textcharactersbytespieces under an English vocabularywhat happened
1the cat772English. One byte per character, and the words are common enough that the vocabulary has them whole.
2le chat773French in plain Latin letters. Same bytes, slightly more pieces, because the corpus contained less French.
3the Russian for cat364Three Cyrillic characters, two bytes each, and the vocabulary has learned few Cyrillic merges, so six bytes become four pieces.
4the Hindi for cat61812Six Devanagari characters at three bytes each, almost no learned merges, so the result sits close to the byte floor. Six times the English piece count for the same word.
4 steps
The character counts are similar across all four rows. The byte counts are not, and the piece counts are worse still, because the number of learned merges falls away exactly where the byte count rises.

This is the step that makes the earlier guarantee true rather than

approximately true, and it is why current tokenizers are described as byte

level. The merging algorithm is unchanged. It simply starts from 256 symbols

rather than from whatever characters happened to appear.

The lesson stops here

4 more paragraphs to go

You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.

The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.

See the planThe contents

This is the reading half

Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.

The contents

The rest of this course

  1. 01Two Obvious Vocabularies, Both of Which Fail
  2. 02Count the Pairs, Join the Winner, Repeatopening only
  3. 03Working From Bytes Makes the Vocabulary Totalyou are here
  4. 04Two Bills That Move in Opposite Directionsopening only
  5. 05The Space Belongs to the Word That Follows Itopening only
  6. 06Where the Digits Get Cut Decides Whether the Sum Worksopening only
  7. 07The Same Sentence, Nine Times the Billopening only
  8. 08The One Decision That Is Made Before Anything Elseopening only

Read alongside