What Ended Up In The Table
Last timeA Row Each
Training arranges the rows so that pieces used in similar places sit near each other, which makes some attributes line up as directions and imports the corpus biases wholesale.
The rows start as noise. Training moves them. The question this lesson answers is
where they end up, and the answer is not something anyone designed, so it has to
be found out by looking.
Start with the only force acting on a row. A row is adjusted when its piece
appears and the model would have predicted the surroundings better with the row
somewhere else. That is the whole mechanism. So the position a row settles into is
determined by what tends to appear around its piece.
Two pieces that appear in similar company are therefore pushed towards similar
positions, because similar positions serve similar surroundings equally well. This
single fact explains almost everything observed about these tables, including the
parts people find surprising.
- the sum of the products of corresponding numbers in the two rows
- the length of row a, which tends to grow with how often its piece appears
- the cosine of the angle between the two rows, from one down to minus one
Attributes as steps
Some differences between rows turn out to be consistent. If the text systematically
changes the surroundings of a word when some attribute changes, then pairs
differing in that attribute end up separated by a roughly similar step, and taking
that step from a new row lands near the corresponding answer.
This is the source of the famous arithmetic with word vectors, and it is worth
being precise about how well it works. The effect is real and measurable. It is
also much weaker than the examples suggest: the nearest row to the computed point
is frequently one of the inputs, which is why demonstrations quietly exclude them,
and accuracy falls off sharply for attributes the text marked inconsistently.
| step | step | what is done | what to watch for |
|---|---|---|---|
| 1 | pick a pair that differ in one attribute | subtract one row from the other | the difference is only consistent if the |
| 2 | add that difference to a third row | arithmetic on numbers, nothing more | the result is a point in the space, not |
| 3 | find the nearest rows to the result | a similarity comparison against every ro | the inputs themselves are usually among |
| 4 | judge the answer | read the nearest remaining row | frequent pieces win disproportionately b |
| 5 | try the same step on other pairs | repeat with many examples | consistency across pairs is the only evi |
No column means anything
It is tempting to imagine that one of the numbers in a row stands for something,
so that column 304 might be how plural a piece is. Nothing in training favours
one set of axes over another. The same arrangement of rows, rotated as a whole,
serves exactly as well, and the procedure has no reason to prefer the unrotated
version.
So the structure lives in relative positions. Comparing rows is informative;
reading a single number out of a row is not. Any claim to have found a meaningful
individual coordinate in a trained table should be met with the question of why
that axis rather than any rotation of it.
The lesson stops here
2 more paragraphs to go
You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.
The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.
See the planThe contentsThis is the reading half
Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.
The contents