ContentsThe library

How Text Becomes Numbers

What Ended Up In The Table

Last timeA Row Each

Training arranges the rows so that pieces used in similar places sit near each other, which makes some attributes line up as directions and imports the corpus biases wholesale.

The rows start as noise. Training moves them. The question this lesson answers is

where they end up, and the answer is not something anyone designed, so it has to

be found out by looking.

Start with the only force acting on a row. A row is adjusted when its piece

appears and the model would have predicted the surroundings better with the row

somewhere else. That is the whole mechanism. So the position a row settles into is

determined by what tends to appear around its piece.

Two pieces that appear in similar company are therefore pushed towards similar

positions, because similar positions serve similar surroundings equally well. This

single fact explains almost everything observed about these tables, including the

parts people find surprising.

FIG 1Why similar company produces similar rows
The last step is where intuition and mechanism part company. If you ask what two nearby rows have in common, the honest answer is that one could be swapped for the other in a sentence without the sentence becoming strange. That often coincides with similar meaning, and sometimes coincides with opposite meaning, and the table has no way to tell those apart.
FIG 2How nearness is measured
the sum of the products of corresponding numbers in the two rows
the length of row a, which tends to grow with how often its piece appears
the cosine of the angle between the two rows, from one down to minus one
Dividing by both lengths is the important part. Without it, frequent pieces would appear similar to everything simply by having longer rows. With it, only direction counts, which is why this measure and not plain distance is what every retrieval system and every nearest-neighbour lookup uses.
FIG 3Reading a similarity value
where most genuinely related pairs land-0.50.5-1exactly opposite directions0unrelated0.3loosely related in practice0.8closely interchangeable1identical directioncosine similarity between two rows
In a real table the useful range is narrow and shifted. Values below about 0.2 are mostly noise, values above 0.9 usually mean two spellings of the same thing, and the interesting comparisons live in between. The negative half of the scale is almost empty, because opposites share company rather than opposing it.

Attributes as steps

Some differences between rows turn out to be consistent. If the text systematically

changes the surroundings of a word when some attribute changes, then pairs

differing in that attribute end up separated by a roughly similar step, and taking

that step from a new row lands near the corresponding answer.

This is the source of the famous arithmetic with word vectors, and it is worth

being precise about how well it works. The effect is real and measurable. It is

also much weaker than the examples suggest: the nearest row to the computed point

is frequently one of the inputs, which is why demonstrations quietly exclude them,

and accuracy falls off sharply for attributes the text marked inconsistently.

FIG 4Taking a step and checking the result
stepstepwhat is donewhat to watch for
1pick a pair that differ in one attributesubtract one row from the otherthe difference is only consistent if the
2add that difference to a third rowarithmetic on numbers, nothing morethe result is a point in the space, not
3find the nearest rows to the resulta similarity comparison against every rothe inputs themselves are usually among
4judge the answerread the nearest remaining rowfrequent pieces win disproportionately b
5try the same step on other pairsrepeat with many examplesconsistency across pairs is the only evi
5 steps
The third and fourth rows are where most demonstrations hide their work. Excluding the inputs is a defensible convention but it is a convention, and it means the raw result of the arithmetic is less impressive than the reported result. The fifth row is the actual test, and it is the one rarely shown.

No column means anything

It is tempting to imagine that one of the numbers in a row stands for something,

so that column 304 might be how plural a piece is. Nothing in training favours

one set of axes over another. The same arrangement of rows, rotated as a whole,

serves exactly as well, and the procedure has no reason to prefer the unrotated

version.

So the structure lives in relative positions. Comparing rows is informative;

reading a single number out of a row is not. Any claim to have found a meaningful

individual coordinate in a trained table should be met with the question of why

that axis rather than any rotation of it.

The lesson stops here

2 more paragraphs to go

You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.

The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.

See the planThe contents

This is the reading half

Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.

The contents

The rest of this course

  1. 01The Layer Nobody Looks At
  2. 02Three Places You Could Cutopening only
  3. 03Nobody Wrote This Listopening only
  4. 04The Number Somebody Had To Chooseopening only
  5. 05One Row, Looked Upopening only
  6. 06What Ended Up In The Tableyou are here
  7. 07The Row Is Only The Starting Pointopening only
  8. 08Complaints That Are Really About This Layeropening only

Read alongside