ContentsThe library

Judging a System, Not a Model

Run Both on the Same Cases and Most of the Noise Disappears

Last timeIs That a Real Improvement

The single technique that buys the most in evaluation: compare per case rather than comparing two averages, and the hard cases stop counting against you.

Why pairing works

The previous lesson left a discouraging number: four hundred cases, six points

before anything can be claimed. Most real improvements are two or three points.

The way out is not more cases. It is to notice where the noise comes from.

When you score version A on one sample and version B on another, the difference

between the two numbers contains two things: the real difference between the

versions, and the difference between the two samples. The second part is

usually much larger. Some cases are simply harder than others, and whichever

version happened to draw the harder sample looks worse.

Run both versions on the same cases and that part is shared. Compare within

each case, and it cancels.

FIG 1What is left after pairing
the uncertainty on the difference, which is the thing you actually want small
the uncertainty on the first version score
the uncertainty on the second version score
how strongly the two versions agree case by case: near 1 when they are similar systems, near 0 when they are unrelated
The last term is the gift. With independent sets it is zero and the errors simply add. With the same cases it is large and positive, and it subtracts. Two versions that agree on nine cases in ten have a correlation near 0.9, which cuts the uncertainty on the difference by about two thirds.

The part worth sitting with is that the benefit grows as the two versions

become more similar. The situation where independent sampling is most hopeless,

a small prompt change that affects a tenth of answers, is precisely where

pairing helps most.

FIG 2Detectable difference against how similar the two versions are
0.002.004.006.008.000.00.20.50.70.9agreement between the two versions, case by case
400 cases, paired400 cases, independent sets
The flat line is what the previous lesson computed. The falling curve is the same four hundred cases used properly. At the agreement levels typical of two versions of one product, the paired design detects two to three points where the independent one needs seven.

Reading a win rate

In practice the comparison is not usually two scores subtracted. It is a

verdict per case: A better, B better, or no difference. The result is a win

rate.

The lesson stops here

4 more paragraphs to go

You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.

The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.

See the planThe contents

This is the reading half

Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.

The contents

The rest of this course

  1. 01The Boundary Decides What the Numbers Mean
  2. 02The Set Is a Sample, and You Chose the Samplingopening only
  3. 03If Two People Score It Differently, You Have No Measurementopening only
  4. 04A Marker That Never Gets Tired and Has Four Known Habitsopening only
  5. 05Three Points on Four Hundred Cases Is Nothingopening only
  6. 06Run Both on the Same Cases and Most of the Noise Disappearsyou are here
  7. 07A Small Fast Set That Blocks, and Two Slower Ones That Do Notopening only
  8. 08Four Reasons Your Number Disagrees With Your Usersopening only

Read alongside