The Boundary Decides What the Numbers Mean
An evaluation measures whatever sits inside the line you drew, and most teams never draw it. Here is where the line goes and what has to be pinned once it is drawn.
Drawing the line
Here is a question that sounds like bureaucracy and decides everything: what,
exactly, are you measuring?
The honest answer for most teams is that nobody wrote it down. There is an
evaluation, it produces a number, the number went from 71 to 74, and a meeting
is now trying to work out whether that is good. It cannot be settled, because
the number is a measurement of something nobody specified.
So draw the line. There is no correct place for it, which is exactly why it has
to be stated rather than assumed. A sensible default is to put everything your
team deploys inside the line, and everything you merely call outside it, but
the defensible version is whichever one you write down and hold to.
Two boundaries that are both reasonable and give very different numbers: the
model alone with a fixed prompt and fixed documents, which tells you about the
model; and the whole chain with live retrieval, which tells you what users got
this morning. The first is reproducible and does not predict the product. The
second predicts the product and cannot be compared across weeks. Most teams
want both and build one by accident.
Everything that can move
Once the line is drawn, the question becomes what inside it can move. The list
is longer than people expect.
The retrieval index is the usual culprit and the least visible. Documents get
added, an embedding model gets upgraded, a filter changes, somebody reindexes
with a different chunk size. None of that is a deploy, none of it appears in
version control, and all of it changes what the model is shown and therefore
what it says.
Prompts are second because they are text, and text does not get reviewed the way
code does. A sentence added to stop one bad behaviour routinely changes the
length, tone and structure of every answer.
Pinning the outside
For each component, there is a decision: frozen for the duration of the test, or
allowed to vary. Anything you do not decide about is varying.
| changes without a deploy | recorded with each resul | pinned during a run | |
|---|---|---|---|
| retrieval index | 1 | 0 | 1 |
| prompt text | 1 | 1 | 1 |
| model version | 0 | 1 | 1 |
| external tool | 1 | 0 | 0 |
| post-processing | 0 | 1 | 1 |
Pinning a live tool usually means recording its responses once and replaying
them. This feels like cheating and is not: you are testing your system, and the
weather service is not your system. If you want to know how you behave when the
weather service misbehaves, that is a different test with deliberately broken
replies in it, which is a better test than hoping the real one fails during a
run.
Recording the version of everything
The last piece is provenance, and it is the cheapest thing in this lesson to
build and the most expensive to retrofit.
| step | model | prompt | index size | score | what happened |
|---|---|---|---|---|---|
| 1 | v4.1 | p-88 | 412000 | 0.71 | The baseline, recorded in full alongside the number. |
| 2 | v4.1 | p-88 | 455000 | 0.64 | A drop of seven points. The model did not change and neither did the prompt, so the index is the only candidate left, and it grew by forty thousand documents. |
| 3 | v4.1 | p-88 | 412000 | 0.7 | Re-run against the pinned older index. The score comes back, which identifies the cause in one run rather than a week. |
Store, with every result: the model and its version, a hash of the prompt text,
a snapshot identifier for the retrieval index, the version of the serving code,
and whether tool responses were live or replayed. That is five fields. They turn
a regression hunt from a week of guessing into a single comparison.
The rest of this course builds on this boundary. The next lesson fills it with
cases drawn from real traffic, and the one after that decides what counts as a
right answer, which is where most evaluations quietly stop meaning anything.
Recap
- Write down the boundary of the thing under test before writing a single case, because every number you produce is a statement about what is inside it.
- Most of what changes a result is not the model: retrieval contents, prompt text, tool responses and post-processing all move scores without any model changing.
- Record the version of every component with each result, or you will have a score you cannot attribute and a regression you cannot locate.
This is the reading half
Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.
The contents