ContentsThe library

Judging a System, Not a Model

The Boundary Decides What the Numbers Mean

An evaluation measures whatever sits inside the line you drew, and most teams never draw it. Here is where the line goes and what has to be pinned once it is drawn.

Drawing the line

Here is a question that sounds like bureaucracy and decides everything: what,

exactly, are you measuring?

The honest answer for most teams is that nobody wrote it down. There is an

evaluation, it produces a number, the number went from 71 to 74, and a meeting

is now trying to work out whether that is good. It cannot be settled, because

the number is a measurement of something nobody specified.

FIG 1A request on its way through
The benchmark measures one box. The user experiences the whole chain. Every box in it can change the final answer, and five of the six change more often than the model does.

So draw the line. There is no correct place for it, which is exactly why it has

to be stated rather than assumed. A sensible default is to put everything your

team deploys inside the line, and everything you merely call outside it, but

the defensible version is whichever one you write down and hold to.

Two boundaries that are both reasonable and give very different numbers: the

model alone with a fixed prompt and fixed documents, which tells you about the

model; and the whole chain with live retrieval, which tells you what users got

this morning. The first is reproducible and does not predict the product. The

second predicts the product and cannot be compared across weeks. Most teams

want both and build one by accident.

Everything that can move

Once the line is drawn, the question becomes what inside it can move. The list

is longer than people expect.

FIG 2Where a hundred regressions actually came from
Indicative proportions from production incident reviews on systems of this shape. The part the benchmark measures is the fourth largest cause. An evaluation that only varies the model is watching the wrong twelve per cent.

The retrieval index is the usual culprit and the least visible. Documents get

added, an embedding model gets upgraded, a filter changes, somebody reindexes

with a different chunk size. None of that is a deploy, none of it appears in

version control, and all of it changes what the model is shown and therefore

what it says.

Prompts are second because they are text, and text does not get reviewed the way

code does. A sentence added to stop one bad behaviour routinely changes the

length, tone and structure of every answer.

Pinning the outside

For each component, there is a decision: frozen for the duration of the test, or

allowed to vary. Anything you do not decide about is varying.

FIG 3A pinning sheet for one system
changes without a deployrecorded with each resulpinned during a run
retrieval index101
prompt text111
model version011
external tool100
post-processing011
One row per component, one column per question, with 1 for yes. The marked cells are the holes: an index that changes silently and is not recorded, and a live external tool that is neither recorded nor pinned. Those two holes are enough to make a month of results incomparable.

Pinning a live tool usually means recording its responses once and replaying

them. This feels like cheating and is not: you are testing your system, and the

weather service is not your system. If you want to know how you behave when the

weather service misbehaves, that is a different test with deliberately broken

replies in it, which is a better test than hoping the real one fails during a

run.

Recording the version of everything

The last piece is provenance, and it is the cheapest thing in this lesson to

build and the most expensive to retrofit.

FIG 4The same question, two weeks apart
stepmodelpromptindex sizescorewhat happened
1v4.1p-884120000.71The baseline, recorded in full alongside the number.
2v4.1p-884550000.64A drop of seven points. The model did not change and neither did the prompt, so the index is the only candidate left, and it grew by forty thousand documents.
3v4.1p-884120000.7Re-run against the pinned older index. The score comes back, which identifies the cause in one run rather than a week.
3 steps
This diagnosis is only possible because the third column was recorded. Without it the sequence is a drop from 71 to 64 with no attached information, and the investigation starts with the model, which is innocent.

Store, with every result: the model and its version, a hash of the prompt text,

a snapshot identifier for the retrieval index, the version of the serving code,

and whether tool responses were live or replayed. That is five fields. They turn

a regression hunt from a week of guessing into a single comparison.

The rest of this course builds on this boundary. The next lesson fills it with

cases drawn from real traffic, and the one after that decides what counts as a

right answer, which is where most evaluations quietly stop meaning anything.

Recap

  • Write down the boundary of the thing under test before writing a single case, because every number you produce is a statement about what is inside it.
  • Most of what changes a result is not the model: retrieval contents, prompt text, tool responses and post-processing all move scores without any model changing.
  • Record the version of every component with each result, or you will have a score you cannot attribute and a regression you cannot locate.

This is the reading half

Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.

The contents

NextChoosing the Cases →

The rest of this course

  1. 01The Boundary Decides What the Numbers Meanyou are here
  2. 02The Set Is a Sample, and You Chose the Samplingopening only
  3. 03If Two People Score It Differently, You Have No Measurementopening only
  4. 04A Marker That Never Gets Tired and Has Four Known Habitsopening only
  5. 05Three Points on Four Hundred Cases Is Nothingopening only
  6. 06Run Both on the Same Cases and Most of the Noise Disappearsopening only
  7. 07A Small Fast Set That Blocks, and Two Slower Ones That Do Notopening only
  8. 08Four Reasons Your Number Disagrees With Your Usersopening only

Read alongside