ContentsThe library

Finding Things by Meaning

The Step That Decides Everything After It

A retrieval system does not answer anything. It hands over a short list of candidates, and the quality of that list is a ceiling on everything that happens next.

A retrieval system is given a question and a collection, and it returns a handful

of pieces of that collection in order. That is all it does. It does not decide

whether any of those pieces answers the question, and it has no notion of whether

they are true.

Why anything needs it

A model holds what it saw during training. Your contract, your incident log and

yesterday's price list are not in there, and retraining to add them is absurd. The

alternative is to put the material in the prompt, which works beautifully until

the material is larger than a prompt can hold or expensive enough that paying to

send all of it on every question stops making sense.

So a smaller thing is put in front: something that can look at the whole

collection cheaply and nominate the few pieces worth paying attention to.

That framing is worth holding on to, because it tells you what the system is

allowed to be bad at. It may rank the right piece second instead of first. It may

include two pieces that turn out to be useless. What it may not do is fail to

nominate the one piece that mattered, because nothing downstream has any way of

asking for a second opinion.

FIG 1Where the step sits, and what each part is responsible for
The division of labour is the whole design. One stage is cheap and must not miss; the next is expensive and must be selective. Systems that try to make a single stage do both jobs are the ones that disappoint.

The two ways it can fail

A shortlist can be wrong in two directions, and the two are not comparable in

seriousness.

If the piece that contains the answer is not in the list, the question cannot be

answered from the material supplied. Nothing further along has any route to it. A

better reader does not help, a better prompt does not help, and a larger model

does not help. The information simply is not present.

If an irrelevant piece is in the list, the damage is smaller and recoverable. A

later stage can push it down, and the reader can ignore it. Not costless, since

irrelevant material takes space and attention, but recoverable in a way the other

failure is not.

This asymmetry has a direct consequence for how the two stages should be tuned,

and it is the opposite of what most people's instinct suggests. The instinct is

to make the first stage precise, so that only good material gets through. The

correct setting is the reverse: make the first stage generous, accept that it

will hand over some rubbish, and spend the precision budget on the stage that

comes after it, where the list is short enough to afford careful work.

FIG 2The ceiling the first stage sets
the fraction of questions whose answer appears somewhere in the shortlist
the fraction of those where the reader actually finds and uses it
the fraction of questions answered correctly, which cannot exceed the product
If the right material reaches the reader six times in ten, then six in ten is the best possible answer rate, and every improvement to the reader is competing for a share of that six. This is why work on prompts sometimes produces nothing measurable.

How long the shortlist should be

Returning more pieces raises the chance the right one is among them. It also

makes that piece a smaller share of what the reader sees, and readers are

measurably worse at using material buried in the middle of a long list than

material near its ends.

FIG 3Two effects of a longer shortlist, pulling opposite ways
0.000.300.600.901.201.015.830.545.360.0pieces returned in the shortlist
the right piece is somewhere in the listthe answer actually uses it
The upper curve keeps rising and flattens. The lower one, which is what you are actually paid for, rises and then falls. The peak sits at a list of around ten, which is close to where production systems converge by trial and error.
FIG 4How often the answer is present, as the shortlist grows
answer present in the shquestion answered correc
one piece returned4138
five6861
twenty8472
one hundred9378
The marked row is the warning. Going from twenty pieces to a hundred adds nine points to what was found and six to what was answered, while multiplying the cost of every question by five. The gap between the two columns is the second stage's problem, not the first stage's.

Which stage failed

The practical consequence of all this is a diagnostic habit. When an answer is

wrong, there are three distinct explanations, and they call for completely

different work.

FIG 5A hundred wrong answers, sorted by where they went wrong
Only the last slice is about the model. The first is about how documents were cut up and indexed, and the second is about ordering. Teams that do not make this split spend their effort on the smallest slice, because it is the one that is easiest to look at.

Making that split requires only one thing: for each failed question, search the

shortlist for the text that would have answered it. If it is absent, the problem

is upstream, and the next three lessons are about the upstream. If it is present,

the problem is ordering or reading, which the later lessons cover.

Doing this properly needs a modest amount of preparation. You need a set of

questions, and for each one, a note saying which piece of the collection contains

the answer. A hundred such questions, written by hand in an afternoon, is enough

to tell you which of the three slices your system is losing to, and that is more

useful than any amount of reading through outputs and forming impressions. The

last lesson of this course is about building and using exactly that set.

It is worth saying plainly that this check is cheap and almost nobody does it.

The usual approach is to look at bad answers, conclude the model is weak, and go

and adjust the instructions. On most systems that is the wrong three quarters of

the problem.

Recap

  • Retrieval returns a ranked shortlist of pieces of a collection, and the model that reads them can only be as good as what it was given, so retrieval quality is a ceiling rather than a contribution.
  • The two ways retrieval can fail are not symmetrical: a piece that was never returned cannot be recovered later, while an irrelevant one that was returned can be filtered out by a further stage.
  • The size of the shortlist is a real decision, because a longer list is more likely to contain the right piece and less likely to have that piece noticed.

This is the reading half

Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.

The contents

NextCutting Documents Into Pieces →

The rest of this course

  1. 01The Step That Decides Everything After Ityou are here
  2. 02The Decision Made Before Anything Is Searchedopening only
  3. 03Neither Side of the Comparison Is Textopening only
  4. 04The Two Methods Fail on Opposite Questionsopening only
  5. 05You Cannot Compare the Question Against Everythingopening only
  6. 06One Knob, and What It Is Really Selling Youopening only
  7. 07The Stage That Can Afford to Be Slowopening only
  8. 08A Hundred Questions You Wrote Yourselfopening only

Read alongside