ContentsThe library

Reading a Benchmark Honestly

Say the Number, the Range, the Conditions and the Limits

Last timeThe Same Model, Many Different Scores

Everything in this course turns into a short list of things a report has to contain, and a shorter list of claims it is then entitled to make. Neither list is expensive to satisfy.

Everything in the previous seven lessons reduces to a short list. A score is a

fraction of a particular set of questions, it carries sampling error, it moves

under formatting, it inflates when the set is picked over, and it collapses when

the answers were in the training data. The practical consequence is a list of

things to write down, and the list is shorter than most people expect.

FIG 1The smallest honest statement about two systems
the claim: how much better one system is than the other, on these questions
both measured on the same question set, so the questions they both get right and both get wrong cancel
the uncertainty of the difference itself, computed from the disagreements and not from the two scores separately
two standard errors, which is the conventional width for a range you expect to contain the truth
This is the whole claim. It names two systems, one question set, a difference and a range. It does not mention a rank, it does not say state of the art, and if the range includes zero it does not get written at all.

The minimum

Alongside it go three more facts, none of which takes a sentence: how many

questions, what the conditions were, and how many configurations were tried

before this one was reported. That is the minimum report. Four items, one of

which is a square root of numbers you already have.

FIG 2How often each ingredient actually appears, as a percentage of published results surveyed
earliermiddlerecent
how many questions were 627178
the extraction and marki212938
an uncertainty range or 182431
the exact text sent to t91422
any check for leakage4711
Everything is improving and everything is low. The marked cell is the worst case, a leakage check appearing in roughly one result in twenty, in a period when training sets were growing fast enough to make leakage likelier every year. The top row is the cheapest item on the list and still misses a fifth of reports.

Ranks are not results

A leaderboard position is a statement about everyone else. It changes when a

system you have never run is added, when another team retunes, when the test set

is refreshed. A paired difference does not. It is a statement about two systems

on a fixed set of questions, and it stays true when the table around it moves.

The lesson stops here

5 more paragraphs to go

You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.

The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.

See the planThe contents

This is the reading half

Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.

The contents

The rest of this course

  1. 01The Number Describes the Questions as Much as the Model
  2. 02How Far the Number Could Have Landed From Where It Didopening only
  3. 03Only the Questions They Disagree About Carry Any Informationopening only
  4. 04The Top of the Table Is Not Where the Best System Isopening only
  5. 05A Model That Has Seen the Exam Is Not Being Examinedopening only
  6. 06The Name on the Box Is a Claim, Not a Descriptionopening only
  7. 07Everything Between the Question and the Numberopening only
  8. 08Say the Number, the Range, the Conditions and the Limitsyou are here

Read alongside