ContentsThe library

How Text Becomes Numbers

Nobody Wrote This List

Last timeWhere to Cut a Word

The list of pieces is built by a loop that repeatedly glues together the commonest adjacent pair, so what counts as one piece is decided entirely by a corpus.

The previous lesson needed a list of pieces where common words are whole and rare

words decompose into common fragments. Nobody writes such a list. It is built by

a loop short enough to state in three sentences, and the loop was borrowed from a

data compression method from the nineteen nineties.

Start by treating every character as its own piece, so the initial list is the

alphabet of the corpus, a few hundred entries. Every possible string is

representable at this point, which is a property worth noting because the

procedure never destroys it.

Then repeat one step. Count every adjacent pair of pieces across the whole

corpus. Find the most frequent pair. Add the two glued together as a new piece,

and rewrite the corpus so that pair is now a single piece everywhere it appears.

Stop when the list reaches the size you decided on in advance.

FIG 1The loop that builds the list
Nothing in this loop knows anything about language. It does not know what a word is, what a prefix is, or that two spellings mean the same thing. It counts adjacent pairs, which is why the output reflects the corpus exactly and nothing else.

Watching it run

The first merges are always the most common letter pairs in the language. After a

few hundred rounds, common short words have been assembled whole. After a few

thousand, most ordinary words are single pieces, and after that the loop spends

its budget on longer multi-word sequences and on rarer vocabulary.

FIG 2The first few rounds on English text
steproundwhat gets mergedwhy it was the most frequent pair
1earlytwo very common lettersletter pairs dominate before anything lo
2a few dozen incommon endings and short function wordsthese sequences repeat in nearly every s
3a few hundred inordinary short words, completetheir letters have already been merged i
4a few thousand inmost common words, each one pieceeach appears often enough to beat any re
5ten thousand and beyondrarer words, inflected forms, and some mall the frequent material has been used
5 steps
The sequence is driven entirely by counting, yet it produces something that looks linguistically sensible for the first few thousand rounds. That resemblance is a coincidence of frequency and not a design goal, and it breaks down exactly where frequency and meaning stop agreeing.
FIG 3The whole procedure
python
pieces = list(characters_in(corpus))
while len(pieces) < target_size:
    pair = most_frequent_adjacent_pair(corpus)
    merges.append(pair)
    pieces.append(join(pair))
    corpus = rewrite_with(corpus, pair)
Six lines, and the finished artefact is the merges list rather than the pieces list. Encoding a new string replays those merges in the same order, which is what makes the cutting of any given string fixed and reproducible. Two tokenisers with identical piece sets but different merge orders will cut the same string differently.

One more property is worth drawing out before moving on. The merges are ordered,

and that order is part of the finished artefact. Encoding a new string does not

search for the best possible cutting; it starts from characters and replays the

saved merges one by one, applying each wherever it fits. The result is whatever

that replay produces. This is cheap, it is deterministic, and it means two

tokenisers holding exactly the same set of pieces in a different order will cut

the same string into different pieces. The list of pieces is not the scheme. The

list of merges, in order, is the scheme.

The lesson stops here

3 more paragraphs to go

You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.

The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.

See the planThe contents

This is the reading half

Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.

The contents

The rest of this course

  1. 01The Layer Nobody Looks At
  2. 02Three Places You Could Cutopening only
  3. 03Nobody Wrote This Listyou are here
  4. 04The Number Somebody Had To Chooseopening only
  5. 05One Row, Looked Upopening only
  6. 06What Ended Up In The Tableopening only
  7. 07The Row Is Only The Starting Pointopening only
  8. 08Complaints That Are Really About This Layeropening only

Read alongside