ContentsThe library

Judging a System, Not a Model

The Set Is a Sample, and You Chose the Sampling

Last timeWhat You Are Actually Testing

An evaluation set is a claim about which inputs matter. Build it from your own traffic, state the sampling behind each group, and keep part of it unseen.

Taking it from real traffic

The temptation at the start is to download something. A public set exists, it

has thousands of cases and a leaderboard, and it can be running by lunchtime.

The problem is that it measures somebody else's distribution of problems. Your

users ask shorter questions, or longer ones, about a domain the public set does

not contain, in a register it does not contain, with the typos and half-finished

sentences that real people produce and curated sets do not. A system can gain

four points on the public set and lose on yours, and the public set will never

tell you.

So the set comes from logs. The work is in how it comes.

FIG 1From a logged request to a case in the set
Five steps, of which the fourth is the one worth insisting on. A set whose cases carry no note of why they were chosen degrades into a pile that somebody eventually deletes because nobody can say what it covers.

Covering the behaviours

Now the central decision. If you sample uniformly from traffic, your set looks

like your traffic, which sounds obviously right and is usually wrong.

Consider a system where 70 per cent of requests are one simple intent handled

perfectly by everything, and 2 per cent are a hard intent where the versions

you are choosing between actually differ. A uniform sample of 300 cases spends

210 on the easy intent and 6 on the hard one. Six cases cannot distinguish

anything, so the evaluation reports no difference, which is false.

FIG 2Traffic share against set share
per cent of trafficper cent of the setper cent of cases where
simple lookup70.020.02.1
multi-step question14.020.04.3
ambiguous request9.020.06.8
out of scope5.020.011.2
adversarial2.020.024.5
Equal quotas rather than proportional ones. The marked row is the argument: adversarial inputs are two per cent of traffic and the versions disagree on a quarter of them, so under proportional sampling almost all of the information in the comparison sits in the handful of cases proportional sampling discards.

Over-sampling a rare group does not distort the result as long as you report by

group as well as in aggregate, and weight the aggregate back to the real

proportions when you want a product-level number. Two numbers from one run: the

weighted total, which estimates the user experience, and the per-group scores,

which say where a change landed.

The lesson stops here

4 more paragraphs to go

You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.

The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.

See the planThe contents

This is the reading half

Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.

The contents

The rest of this course

  1. 01The Boundary Decides What the Numbers Mean
  2. 02The Set Is a Sample, and You Chose the Samplingyou are here
  3. 03If Two People Score It Differently, You Have No Measurementopening only
  4. 04A Marker That Never Gets Tired and Has Four Known Habitsopening only
  5. 05Three Points on Four Hundred Cases Is Nothingopening only
  6. 06Run Both on the Same Cases and Most of the Noise Disappearsopening only
  7. 07A Small Fast Set That Blocks, and Two Slower Ones That Do Notopening only
  8. 08Four Reasons Your Number Disagrees With Your Usersopening only

Read alongside