Where Training Data Comes From
The Benchmark That Lies in the Flattering Direction
Last timeHow Much of Each
If the evaluation set leaked into training, every number you report is wrong upwards, which is the direction nobody questions. Measuring the overlap is the only defence.
Nobody does this on purpose
Contamination is the presence of evaluation material in the training corpus,
and the first thing to understand is that it is not cheating. Nobody copies a
test file into the training data.
It arrives by routes that are entirely ordinary.
The uncomfortable corollary is that the more useful and widely used a benchmark
is, the more thoroughly it has been written about, and the more contaminated
every subsequent corpus becomes. Benchmarks decay by being popular.
Why it fails upwards
A contaminated model does not merely score higher. It scores higher in a
specific, quantifiable way, because for the leaked portion of the test set it is
recalling rather than reasoning.
The lesson stops here
4 more paragraphs to go
You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.
The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.
See the planThe contentsThis is the reading half
Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.
The contents