The library

Putting a model into service

Be able to look at a reported score, work out roughly how much of it is real, decide whether a gap between two systems means anything, and say what the number does and does not entitle anyone to claim

Reading a Benchmark Honestly

Every claim about a model rests on a number measured on a few hundred questions. This course covers where that number comes from, how much it wobbles, when a gap is real, and the ways a benchmark stops measuring what it says.

8 lessons, written and corrected before you arrived. Reading them here needs no account. The first reads the whole way through; the others open and then stop, because a page nobody owns cannot tell who is reading it. Starting the course gives you your own copy, where every idea has problems standing under it and you can ask about any sentence.

Start reading

  1. 01The Number Describes the Questions as Much as the ModelA benchmark score is a fraction of a few hundred questions somebody chose, marked by a rule somebody wrote, under conditions somebody picked. None of those three are the model.
  2. 02How Far the Number Could Have Landed From Where It Didopening onlyA score measured on a few hundred questions carries several points of wobble that come from nothing but which questions were drawn. Here is how to work out how many.
  3. 03Only the Questions They Disagree About Carry Any Informationopening onlyComparing two systems on the same questions is a far sharper measurement than comparing two separate scores, because everything the questions have in common cancels out.
  4. 04The Top of the Table Is Not Where the Best System Isopening onlyTake enough noisy measurements and keep the largest one and you have measured luck. Leaderboards do this by design, and a test set reused for a year stops being a test set.
  5. 05A Model That Has Seen the Exam Is Not Being Examinedopening onlyBenchmark questions are published on the open web, training corpora are scraped from the open web, and the arithmetic of that goes one way. Here is what it does to a score and how to detect it.
  6. 06The Name on the Box Is a Claim, Not a Descriptionopening onlyA benchmark called reasoning measures whatever its questions happen to reward. Shortcuts, wrong answer keys and narrow coverage all open a gap between the name and the measurement.
  7. 07Everything Between the Question and the Numberopening onlyOne model, one question set, a dozen defensible ways of running the evaluation, and a spread of fifteen points. The conditions are not a footnote; they are part of the measurement.
  8. 08Say the Number, the Range, the Conditions and the Limitsopening onlyEverything in this course turns into a short list of things a report has to contain, and a shorter list of claims it is then entitled to make. Neither list is expensive to satisfy.