Neither Side of the Comparison Is Text
Last timeCutting Documents Into Pieces
A question is never matched against a document. Both are turned into lists of numbers by a model trained to put related things near each other, and everything follows from that.
The system has pieces now. What does it do with them?
Not store their words for looking up. It runs each piece through a model that
turns text into a list of a few hundred numbers, and stores that. When a question
arrives, the question goes through a model too, and what gets compared is two
lists of numbers.
The comparison itself
What gets computed is a single number per piece, and the usual choice is the
cosine, which measures the angle between the two lists and ignores how long they
are.
- the list of numbers for the question
- the stored list for one piece
- the sum of the products of corresponding entries
- the length of the question list, which divides out so that only direction matters
It is worth noticing what this choice throws away. Two pieces of identical
subject matter, one a careful page and one a careless sentence, get compared on
direction alone, so nothing about length, authority or completeness enters the
score. That information is not lost from the system, since the text is still
stored, but it is lost from this comparison, and recovering it is a job for a
later stage rather than this one.
The lesson stops here
6 more paragraphs to go
You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.
The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.
See the planThe contentsThis is the reading half
Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.
The contents