The Step That Decides Everything After It
A retrieval system does not answer anything. It hands over a short list of candidates, and the quality of that list is a ceiling on everything that happens next.
A retrieval system is given a question and a collection, and it returns a handful
of pieces of that collection in order. That is all it does. It does not decide
whether any of those pieces answers the question, and it has no notion of whether
they are true.
Why anything needs it
A model holds what it saw during training. Your contract, your incident log and
yesterday's price list are not in there, and retraining to add them is absurd. The
alternative is to put the material in the prompt, which works beautifully until
the material is larger than a prompt can hold or expensive enough that paying to
send all of it on every question stops making sense.
So a smaller thing is put in front: something that can look at the whole
collection cheaply and nominate the few pieces worth paying attention to.
That framing is worth holding on to, because it tells you what the system is
allowed to be bad at. It may rank the right piece second instead of first. It may
include two pieces that turn out to be useless. What it may not do is fail to
nominate the one piece that mattered, because nothing downstream has any way of
asking for a second opinion.
The two ways it can fail
A shortlist can be wrong in two directions, and the two are not comparable in
seriousness.
If the piece that contains the answer is not in the list, the question cannot be
answered from the material supplied. Nothing further along has any route to it. A
better reader does not help, a better prompt does not help, and a larger model
does not help. The information simply is not present.
If an irrelevant piece is in the list, the damage is smaller and recoverable. A
later stage can push it down, and the reader can ignore it. Not costless, since
irrelevant material takes space and attention, but recoverable in a way the other
failure is not.
This asymmetry has a direct consequence for how the two stages should be tuned,
and it is the opposite of what most people's instinct suggests. The instinct is
to make the first stage precise, so that only good material gets through. The
correct setting is the reverse: make the first stage generous, accept that it
will hand over some rubbish, and spend the precision budget on the stage that
comes after it, where the list is short enough to afford careful work.
- the fraction of questions whose answer appears somewhere in the shortlist
- the fraction of those where the reader actually finds and uses it
- the fraction of questions answered correctly, which cannot exceed the product
How long the shortlist should be
Returning more pieces raises the chance the right one is among them. It also
makes that piece a smaller share of what the reader sees, and readers are
measurably worse at using material buried in the middle of a long list than
material near its ends.
| answer present in the sh | question answered correc | |
|---|---|---|
| one piece returned | 41 | 38 |
| five | 68 | 61 |
| twenty | 84 | 72 |
| one hundred | 93 | 78 |
Which stage failed
The practical consequence of all this is a diagnostic habit. When an answer is
wrong, there are three distinct explanations, and they call for completely
different work.
Making that split requires only one thing: for each failed question, search the
shortlist for the text that would have answered it. If it is absent, the problem
is upstream, and the next three lessons are about the upstream. If it is present,
the problem is ordering or reading, which the later lessons cover.
Doing this properly needs a modest amount of preparation. You need a set of
questions, and for each one, a note saying which piece of the collection contains
the answer. A hundred such questions, written by hand in an afternoon, is enough
to tell you which of the three slices your system is losing to, and that is more
useful than any amount of reading through outputs and forming impressions. The
last lesson of this course is about building and using exactly that set.
It is worth saying plainly that this check is cheap and almost nobody does it.
The usual approach is to look at bad answers, conclude the model is weak, and go
and adjust the instructions. On most systems that is the wrong three quarters of
the problem.
Recap
- Retrieval returns a ranked shortlist of pieces of a collection, and the model that reads them can only be as good as what it was given, so retrieval quality is a ceiling rather than a contribution.
- The two ways retrieval can fail are not symmetrical: a piece that was never returned cannot be recovered later, while an irrelevant one that was returned can be filtered out by a further stage.
- The size of the shortlist is a real decision, because a longer list is more likely to contain the right piece and less likely to have that piece noticed.
This is the reading half
Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.
The contents