ContentsThe library

How Text Becomes Numbers

One Row, Looked Up

Last timeHow Long Should the List Be

Every entry in the list owns one row of numbers, the model reads nothing but those rows, and the rows are learned during training rather than assigned.

After cutting, a passage of text is a list of whole numbers. Forty one thousand

two hundred and seven, then nine, then twelve thousand. These are positions in a

list and nothing else. The next step is the one people expect to be complicated

and it is the simplest in the whole path.

There is a table. It has one row for every entry in the list of pieces and one

column for every number in the model width. Identifier nine means row nine. Read

it. That is the entire operation.

FIG 1From identifiers to a block of numbers
Notice what happens at the last two steps. The original text is not carried forward. Anything the cutting lost, including where the spaces were and how an unusual word was broken up, is unrecoverable from this point onwards. The model is working with the rectangle, not with your string.
FIG 2The lookup written as arithmetic
the lookup table, with one row per entry and one column per model dimension
a vector that is one in position i and zero everywhere else
the row belonging to entry i, which is what the model receives
This is how the operation is written in papers, and it is worth being able to read even though it describes something nobody computes. Multiplying by a vector of zeros with a single one in it selects a row. Every implementation skips the multiplication and reads the row directly, which is why this step costs almost nothing despite the table being one of the largest objects in the model.

Nobody writes the numbers

The obvious question is where the values in the table come from. The answer is

that they start as small random numbers and are then changed by training, using

precisely the same procedure that changes every other weight in the model.

When a piece appears in a training example and the model would have done better

had that piece been represented slightly differently, the row moves slightly in

that direction. Repeated billions of times over billions of examples, the rows

drift into an arrangement that makes the rest of the model's job easier. There is

no step at which anybody decides what a row should mean.

FIG 3The lookup and what it costs
python
table = random_small((vocab_size, width))
rows = table[ids]
assert rows.shape == (len(ids), width)
table_bytes = vocab_size * width * 4
Four lines. The first says the table is initialised rather than written. The second is the lookup, which is an index operation and not a matrix product. The third records the shape of what comes out, one row per piece. The fourth is the storage, four bytes per number at ordinary precision, which for fifty thousand entries and a width of 768 is about 150 megabytes before anything is trained.

Two numbers, two decisions

The length of the list and the width of the rows are separate choices. The list

length was the subject of the previous lesson. The width is almost always

inherited from the rest of the model, because the rows are fed straight into the

first layer and have to be the size that layer expects.

That means the table size is a product of one decision made about text and one

decision made about model capacity. Widening the model widens the table even

though nothing about the text changed, which is part of why the table is a smaller

share of a large model than of a small one.

The lesson stops here

3 more paragraphs to go

You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.

The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.

See the planThe contents

This is the reading half

Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.

The contents

The rest of this course

  1. 01The Layer Nobody Looks At
  2. 02Three Places You Could Cutopening only
  3. 03Nobody Wrote This Listopening only
  4. 04The Number Somebody Had To Chooseopening only
  5. 05One Row, Looked Upyou are here
  6. 06What Ended Up In The Tableopening only
  7. 07The Row Is Only The Starting Pointopening only
  8. 08Complaints That Are Really About This Layeropening only

Read alongside