One Row, Looked Up
Last timeHow Long Should the List Be
Every entry in the list owns one row of numbers, the model reads nothing but those rows, and the rows are learned during training rather than assigned.
After cutting, a passage of text is a list of whole numbers. Forty one thousand
two hundred and seven, then nine, then twelve thousand. These are positions in a
list and nothing else. The next step is the one people expect to be complicated
and it is the simplest in the whole path.
There is a table. It has one row for every entry in the list of pieces and one
column for every number in the model width. Identifier nine means row nine. Read
it. That is the entire operation.
- the lookup table, with one row per entry and one column per model dimension
- a vector that is one in position i and zero everywhere else
- the row belonging to entry i, which is what the model receives
Nobody writes the numbers
The obvious question is where the values in the table come from. The answer is
that they start as small random numbers and are then changed by training, using
precisely the same procedure that changes every other weight in the model.
When a piece appears in a training example and the model would have done better
had that piece been represented slightly differently, the row moves slightly in
that direction. Repeated billions of times over billions of examples, the rows
drift into an arrangement that makes the rest of the model's job easier. There is
no step at which anybody decides what a row should mean.
table = random_small((vocab_size, width))
rows = table[ids]
assert rows.shape == (len(ids), width)
table_bytes = vocab_size * width * 4Two numbers, two decisions
The length of the list and the width of the rows are separate choices. The list
length was the subject of the previous lesson. The width is almost always
inherited from the rest of the model, because the rows are fed straight into the
first layer and have to be the size that layer expects.
That means the table size is a product of one decision made about text and one
decision made about model capacity. Widening the model widens the table even
though nothing about the text changed, which is part of why the table is a smaller
share of a large model than of a small one.
The lesson stops here
3 more paragraphs to go
You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.
The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.
See the planThe contentsThis is the reading half
Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.
The contents