ContentsThe library

How Text Becomes Numbers

Three Places You Could Cut

Last timeNothing Reads Letters

Cutting text into characters, into words, or into something in between are three different trades between the length of the list and the length of the sequence.

The cutting step needs a list of known pieces and a rule for matching text

against it. Everything about that step follows from how big the pieces are, and

there are three natural answers.

Cut at every character. Cut at every space. Or cut somewhere in between, at

boundaries chosen by looking at data rather than by any rule a person would

write. The third answer is what everything uses now, and the reason is clearest

after seeing what the other two cost.

Characters

Take the pieces to be single characters. The list is tiny, a few hundred entries

covering every letter, digit and mark anyone will type. Nothing is ever unknown,

because every string is made of characters by definition. Misspellings, names,

code and invented words all go through without special handling, and the round

trip is trivially exact.

Then look at the sequence length. An average English word is around five

characters including the space, so a hundred-word passage becomes five hundred

pieces instead of a hundred. Everything downstream scales with that number, and

the attention mechanism inside a model scales worse than linearly, so a factor of

five in length is considerably more than a factor of five in cost.

There is a second cost that matters more than it looks. The model now has to

spend capacity learning that certain character sequences form words at all, which

is effort not spent on anything else.

FIG 1Where sequence length comes from
the sequence length handed to the model
the number of words in the text
the average number of pieces per word, which the scheme decides
The only lever a cutting scheme has is the middle factor. Characters put it around five for English, words put it at one, and the compromise lands between one and two for ordinary text. Every cost downstream is a function of the result, which is why this single number is worth caring about.

Words

Now take the pieces to be words, split on spaces. Sequences are as short as they

can possibly be, which is the whole attraction.

The list is the problem. A vocabulary that covers ordinary English text runs into

the hundreds of thousands of entries, each needing its own row in the table, and

the table is one of the largest objects in a small model. Worse, that enormous

list is still not enough. Names are effectively unlimited. Typos are unlimited.

New words appear constantly, and any word not on the list has to fall into a

single catch-all entry.

The lesson stops here

6 more paragraphs to go

You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.

The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.

See the planThe contents

This is the reading half

Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.

The contents

The rest of this course

  1. 01The Layer Nobody Looks At
  2. 02Three Places You Could Cutyou are here
  3. 03Nobody Wrote This Listopening only
  4. 04The Number Somebody Had To Chooseopening only
  5. 05One Row, Looked Upopening only
  6. 06What Ended Up In The Tableopening only
  7. 07The Row Is Only The Starting Pointopening only
  8. 08Complaints That Are Really About This Layeropening only

Read alongside