The Set Is a Sample, and You Chose the Sampling
Last timeWhat You Are Actually Testing
An evaluation set is a claim about which inputs matter. Build it from your own traffic, state the sampling behind each group, and keep part of it unseen.
Taking it from real traffic
The temptation at the start is to download something. A public set exists, it
has thousands of cases and a leaderboard, and it can be running by lunchtime.
The problem is that it measures somebody else's distribution of problems. Your
users ask shorter questions, or longer ones, about a domain the public set does
not contain, in a register it does not contain, with the typos and half-finished
sentences that real people produce and curated sets do not. A system can gain
four points on the public set and lose on yours, and the public set will never
tell you.
So the set comes from logs. The work is in how it comes.
Covering the behaviours
Now the central decision. If you sample uniformly from traffic, your set looks
like your traffic, which sounds obviously right and is usually wrong.
Consider a system where 70 per cent of requests are one simple intent handled
perfectly by everything, and 2 per cent are a hard intent where the versions
you are choosing between actually differ. A uniform sample of 300 cases spends
210 on the easy intent and 6 on the hard one. Six cases cannot distinguish
anything, so the evaluation reports no difference, which is false.
| per cent of traffic | per cent of the set | per cent of cases where | |
|---|---|---|---|
| simple lookup | 70.0 | 20.0 | 2.1 |
| multi-step question | 14.0 | 20.0 | 4.3 |
| ambiguous request | 9.0 | 20.0 | 6.8 |
| out of scope | 5.0 | 20.0 | 11.2 |
| adversarial | 2.0 | 20.0 | 24.5 |
Over-sampling a rare group does not distort the result as long as you report by
group as well as in aggregate, and weight the aggregate back to the real
proportions when you want a product-level number. Two numbers from one run: the
weighted total, which estimates the user experience, and the per-group scores,
which say where a change landed.
The lesson stops here
4 more paragraphs to go
You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.
The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.
See the planThe contentsThis is the reading half
Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.
The contents