The One Decision That Is Made Before Anything Else
Last timeWhat It Costs in Other Languages
The vocabulary is fixed before the first training step and everything above it is built on the identity of its entries, so changing it later invalidates the model rather than updating it.
What exactly is frozen
Everything in this course is decided before the first training step: how big
the vocabulary is, whether spaces attach to words, how numbers get cut, which
languages get merges. After that step none of it can be revisited, and it is
worth being exact about why, because the reason is more specific than the usual
statement that the model depends on the tokenizer.
Two things depend on it.
The first is the embedding table: one row of numbers for each piece id. That
row is the model's learned meaning for that id, built up over the whole of
training from every context the piece appeared in. Row 1,189 means whatever
piece 1,189 was.
The second is every weight above the embedding table, which is almost all of
the model. Those weights were fitted to the statistics of text cut this
particular way: which pieces follow which, how long a word takes, where the
boundaries fall inside a number. A model trained on text where common words
arrive whole has learned different regularities than one trained on text
arriving character by character.
What a change invalidates
Now consider editing the merge list after training: adding a few entries,
removing rare ones, refitting the whole thing on a better corpus.
The first thing to understand is that nothing errors. The new list still
produces integers, still inside the model's range, so the model accepts them
and answers. It has simply been handed a different language. Where id 1,189
used to be a common English word it is now some fragment of a different script,
and the row of numbers the model learned for the first meaning is applied to
the second.
| step | id 262 | id 1189 | id 318 | what the model reads | what happened |
|---|---|---|---|---|---|
| 1 | the | cat | sat | the cat sat | Under the list the model was trained on, these three ids are a sentence it has seen thousands of variants of. |
| 2 | th | oriz | tal | th oriz tal | Under a refitted list the same three integers name different pieces. The model is not told. It applies the meanings it learned to pieces that no longer have them. |
The second thing is that this is not repairable by retraining a small part.
Replacing only the embedding rows leaves every weight above them fitted to the
old statistics. Fine-tuning briefly on top gets some of the way back and leaves
a model that is subtly worse at everything, in ways that are hard to attribute
afterwards.
The lesson stops here
3 more paragraphs to go
You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.
The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.
See the planThe contentsThis is the reading half
Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.
The contents