ContentsThe library

Reading a Benchmark Honestly

Everything Between the Question and the Number

Last timeThe Gap Between the Name and the Questions

One model, one question set, a dozen defensible ways of running the evaluation, and a spread of fifteen points. The conditions are not a footnote; they are part of the measurement.

Two groups report a score for the same model on the same public question set.

One says 68, the other says 75. Neither is lying, neither made an error, and

both numbers are correct. The difference is in everything that sits between the

question and the number, which neither report described.

FIG 1The machinery between a question and a score
Seven decisions sit between the file and the number, and a published result typically describes none of them. The gap between two groups reporting the same model is almost always somewhere in this chain rather than in the model or the questions.

The size of the effect

Everyone who runs evaluations knows these choices exist. The usual assumption is

that they are small, a fraction of a point here or there, noise around a stable

underlying ability. The uncomfortable part is that they are not small. They are

larger than the differences most reports are written about, and they are larger

than the sampling error discussed earlier in this course, which at least has the

courtesy of being computable.

FIG 2Three models under four formattings of the same questions
model Amodel Bmodel C
options lettered, instru716874
options numbered, instru647066
options lettered, instru596170
plain list, no instructi667263
Read down any column and the spread is ten points or more from formatting alone. Read across any row and the winner changes: A wins the second row, B the fourth, C the first and third. Whoever picks the format picks the winner, and the choice looks entirely innocent.
FIG 3One model, twenty reasonable formattings of the same set
spread from formatting alonethe gap being argued about55606570758059worst66median74bestscore in percentage points
The shorter span is a difference between two systems that an announcement was built on. The longer one is the same model measured twenty ways. Anyone free to choose the formatting can produce either side of the argued gap without touching the model, which is why a number reported without its conditions carries very little weight.

One at a time

The effects are easiest to see when the settings are changed singly.

The lesson stops here

3 more paragraphs to go

You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.

The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.

See the planThe contents

This is the reading half

Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.

The contents

The rest of this course

  1. 01The Number Describes the Questions as Much as the Model
  2. 02How Far the Number Could Have Landed From Where It Didopening only
  3. 03Only the Questions They Disagree About Carry Any Informationopening only
  4. 04The Top of the Table Is Not Where the Best System Isopening only
  5. 05A Model That Has Seen the Exam Is Not Being Examinedopening only
  6. 06The Name on the Box Is a Claim, Not a Descriptionopening only
  7. 07Everything Between the Question and the Numberyou are here
  8. 08Say the Number, the Range, the Conditions and the Limitsopening only

Read alongside