The Layer Nobody Looks At
A model is a long chain of arithmetic, so text has to arrive as numbers, and the two steps that convert it decide more about behaviour than most people expect.
Open any model and look for the part that handles letters. There is not one. From
the first layer to the last, the operations are multiplications of matrices,
additions, and a handful of simple functions applied to numbers. None of them can
take a character as an argument, because none of them know what a character is.
So before any of it runs, text has to become numbers. That is not a controversial
statement and it is usually passed over in a sentence. What gets missed is that
the conversion is a design with consequences, and a great many complaints about
model behaviour are complaints about this layer wearing a disguise.
The identifiers mean nothing
It is worth being blunt about the middle step because it is routinely
misunderstood. The identifier assigned to a piece is a position in a list. If the
piece for the word large sits at position 4521 and the piece for the word small
sits at 4522, that proximity says nothing whatsoever. They are neighbours in the
list the way two unrelated words are neighbours in a dictionary.
Nothing downstream ever does arithmetic on the identifier. It is used once, to
index a table, and then discarded. This is why you cannot inspect a model by
looking at token numbers, and why renumbering the entire vocabulary, as long as
the table is renumbered to match, changes nothing at all.
ids = tokenizer.encode(text)
rows = embedding_table[ids]
text_again = tokenizer.decode(ids)
assert text_again == textEverything the same width
The table has one row per piece in the vocabulary, and every row has the same
length. That length is a design choice, usually a few hundred to a few thousand,
and it is the same for the most common piece in the language and for the rarest.
This is forced rather than chosen. The first thing the model does is multiply the
incoming rows by a matrix, and a matrix has a fixed number of columns. There is
no mechanism for giving an important piece more numbers than an unimportant one.
What varies is not the width but the content: training arranges for rare pieces to
sit in less useful parts of the space, which is a different thing from giving them
less room.
- the total count of numbers handed to the model for this text
- the number of pieces the text was cut into, which depends on the cutting scheme
- the width of a row, fixed for the whole vocabulary
What the choices actually decide
Here is why this layer deserves a course rather than a paragraph. Four separate
things are settled before the model computes anything.
How long the sequence is, which sets the cost of running the model, because the
work of the attention mechanism grows faster than linearly in the length. Which
spellings look related, since two forms of a word that get cut differently have
no reason to share anything. Whether numbers arrive digit by digit or in clumps,
which is most of the story behind arithmetic errors. And how efficiently different
languages are handled, since a vocabulary built mostly from one language cuts
others into many more pieces.
| step | stage | what exists at this point | who decided it |
|---|---|---|---|
| 1 | typed | a string of characters, exactly as enter | the person |
| 2 | cut | a list of pieces, each one present in th | whoever built the vocabulary |
| 3 | numbered | a list of integer labels, one per piece | the position of each piece in the list |
| 4 | looked up | a row of numbers per piece, all the same | training |
| 5 | consumed | one array of numbers, length times width | the model architecture |
Reversibility
The last requirement is easy to state and easy to get wrong. A model produces
pieces, and those pieces must turn back into text that a person can read, byte for
byte. Not approximately, not after normalisation, exactly.
This rules out a number of conveniences that older text pipelines took for
granted. Lowercasing everything loses information that cannot be recovered.
Collapsing runs of spaces breaks code and poetry. Stripping accents makes several
languages unwritable. Treating the text as words separated by spaces fails for
languages that do not use spaces that way.
| Conversion scheme | List stays small | Sequence stays short | Handles unknown words | Round trip exact |
|---|---|---|---|---|
| Split on spaces into words | no | yes | no | yes |
| Split into single characters | yes | no | yes | yes |
| Split into learned pieces | yes | nearly | yes | yes |
| Normalise, then split into words | no | yes | no | no |
The third row is what everything uses now, and the table shows why: it is the only
scheme with an acceptable answer in all four columns at once. Each of the others is
better in one column and disqualified in another, which is the pattern the next two
lessons examine in detail.
The rest of the course takes these in order. Where to cut and what each choice
costs. How a list of pieces is built from a corpus rather than written by hand.
What making the list longer buys and what it costs. Then the table itself: what is
in a row, what training arranges it to mean, and how the model changes it once
context is available. The course ends on the failures, because once the layer is
understood, several familiar complaints about model behaviour stop being mysteries
and become predictable consequences of where the text was cut.
Recap
- Everything inside a model is multiplication and addition on fixed-width rows of numbers, so text must be converted before anything can happen.
- The conversion is two separate steps: cut the text into pieces from a fixed list, then replace each piece with a row looked up in a table.
- The identifier of a piece is an arbitrary label with no order or size meaning, while the row it points to is where all the meaning lives.
This is the reading half
Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.
The contents