ContentsThe library

How Text Becomes Numbers

The Layer Nobody Looks At

A model is a long chain of arithmetic, so text has to arrive as numbers, and the two steps that convert it decide more about behaviour than most people expect.

Open any model and look for the part that handles letters. There is not one. From

the first layer to the last, the operations are multiplications of matrices,

additions, and a handful of simple functions applied to numbers. None of them can

take a character as an argument, because none of them know what a character is.

So before any of it runs, text has to become numbers. That is not a controversial

statement and it is usually passed over in a sentence. What gets missed is that

the conversion is a design with consequences, and a great many complaints about

model behaviour are complaints about this layer wearing a disguise.

FIG 1The whole path from typing to arithmetic
Two conversions, not one. The first is a lookup of strings in a list and the second is a lookup of rows in a table, and they are built, trained and debugged separately. Most confusion about this layer comes from treating them as a single thing called tokenisation.

The identifiers mean nothing

It is worth being blunt about the middle step because it is routinely

misunderstood. The identifier assigned to a piece is a position in a list. If the

piece for the word large sits at position 4521 and the piece for the word small

sits at 4522, that proximity says nothing whatsoever. They are neighbours in the

list the way two unrelated words are neighbours in a dictionary.

Nothing downstream ever does arithmetic on the identifier. It is used once, to

index a table, and then discarded. This is why you cannot inspect a model by

looking at token numbers, and why renumbering the entire vocabulary, as long as

the table is renumbered to match, changes nothing at all.

FIG 2What the two steps look like from outside
python
ids = tokenizer.encode(text)
rows = embedding_table[ids]
text_again = tokenizer.decode(ids)
assert text_again == text
Three lines and an assertion. The first line is the text algorithm, the second is a table lookup that produces the numbers the model consumes, and the third runs the first backwards. The assertion is the requirement that makes generation possible at all, and schemes that quietly strip accents or collapse whitespace fail it in ways that only show up much later.

Everything the same width

The table has one row per piece in the vocabulary, and every row has the same

length. That length is a design choice, usually a few hundred to a few thousand,

and it is the same for the most common piece in the language and for the rarest.

This is forced rather than chosen. The first thing the model does is multiply the

incoming rows by a matrix, and a matrix has a fixed number of columns. There is

no mechanism for giving an important piece more numbers than an unimportant one.

What varies is not the width but the content: training arranges for rare pieces to

sit in less useful parts of the space, which is a different thing from giving them

less room.

FIG 3How much arrives at the model
the total count of numbers handed to the model for this text
the number of pieces the text was cut into, which depends on the cutting scheme
the width of a row, fixed for the whole vocabulary
Both factors are decisions made before the model runs. The width is fixed when the model is designed. The length depends on the cutting scheme and on the text, and a scheme that cuts a given sentence into twice as many pieces doubles everything downstream. That is the lever the next three lessons are about.

What the choices actually decide

Here is why this layer deserves a course rather than a paragraph. Four separate

things are settled before the model computes anything.

How long the sequence is, which sets the cost of running the model, because the

work of the attention mechanism grows faster than linearly in the length. Which

spellings look related, since two forms of a word that get cut differently have

no reason to share anything. Whether numbers arrive digit by digit or in clumps,

which is most of the story behind arithmetic errors. And how efficiently different

languages are handled, since a vocabulary built mostly from one language cuts

others into many more pieces.

FIG 4One short string through the whole path
stepstagewhat exists at this pointwho decided it
1typeda string of characters, exactly as enterthe person
2cuta list of pieces, each one present in thwhoever built the vocabulary
3numbereda list of integer labels, one per piecethe position of each piece in the list
4looked upa row of numbers per piece, all the sametraining
5consumedone array of numbers, length times widththe model architecture
5 steps
Notice who decides each stage. Only the last two involve the model at all, and the two that set the sequence length were fixed long before by people building a vocabulary on a corpus that may have nothing to do with the text in front of it.

Reversibility

The last requirement is easy to state and easy to get wrong. A model produces

pieces, and those pieces must turn back into text that a person can read, byte for

byte. Not approximately, not after normalisation, exactly.

This rules out a number of conveniences that older text pipelines took for

granted. Lowercasing everything loses information that cannot be recovered.

Collapsing runs of spaces breaks code and poetry. Stripping accents makes several

languages unwritable. Treating the text as words separated by spaces fails for

languages that do not use spaces that way.

FIG 5
Conversion schemeList stays smallSequence stays shortHandles unknown wordsRound trip exact
Split on spaces into wordsnoyesnoyes
Split into single charactersyesnoyesyes
Split into learned piecesyesnearlyyesyes
Normalise, then split into wordsnoyesnono

The third row is what everything uses now, and the table shows why: it is the only

scheme with an acceptable answer in all four columns at once. Each of the others is

better in one column and disqualified in another, which is the pattern the next two

lessons examine in detail.

The rest of the course takes these in order. Where to cut and what each choice

costs. How a list of pieces is built from a corpus rather than written by hand.

What making the list longer buys and what it costs. Then the table itself: what is

in a row, what training arranges it to mean, and how the model changes it once

context is available. The course ends on the failures, because once the layer is

understood, several familiar complaints about model behaviour stop being mysteries

and become predictable consequences of where the text was cut.

Recap

  • Everything inside a model is multiplication and addition on fixed-width rows of numbers, so text must be converted before anything can happen.
  • The conversion is two separate steps: cut the text into pieces from a fixed list, then replace each piece with a row looked up in a table.
  • The identifier of a piece is an arbitrary label with no order or size meaning, while the row it points to is where all the meaning lives.

This is the reading half

Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.

The contents

NextWhere to Cut a Word →

The rest of this course

  1. 01The Layer Nobody Looks Atyou are here
  2. 02Three Places You Could Cutopening only
  3. 03Nobody Wrote This Listopening only
  4. 04The Number Somebody Had To Chooseopening only
  5. 05One Row, Looked Upopening only
  6. 06What Ended Up In The Tableopening only
  7. 07The Row Is Only The Starting Pointopening only
  8. 08Complaints That Are Really About This Layeropening only

Read alongside