A Model That Has Seen the Exam Is Not Being Examined
Last timeWhat Happens on a Leaderboard
Benchmark questions are published on the open web, training corpora are scraped from the open web, and the arithmetic of that goes one way. Here is what it does to a score and how to detect it.
A benchmark is published. It goes on a public repository so that others can use
it. People write tutorials about it, quote example questions in articles,
include it in evaluation libraries, and translate it. A year later, somebody
assembles a training corpus by crawling the web.
Nobody has to misbehave
The more useful a benchmark is, the more widely it is copied, and the more
thoroughly it ends up in corpora. Popularity and contamination are the same
process seen from two sides.
What it does to the number
A model that has seen a question and its answer can produce the answer without
any of the ability the question was written to test. The marking rule is
satisfied. The score goes up. What the number now measures is memory of this
particular set, and because the contaminated items are the popular ones, the
inflation lands exactly where people are looking.
The lesson stops here
5 more paragraphs to go
You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.
The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.
See the planThe contentsThis is the reading half
Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.
The contents