ContentsThe library

Reading a Benchmark Honestly

The Number Describes the Questions as Much as the Model

A benchmark score is a fraction of a few hundred questions somebody chose, marked by a rule somebody wrote, under conditions somebody picked. None of those three are the model.

Somebody tells you a model scores eighty four. Before anything else, it is worth

being precise about what has just been said, because the sentence sounds like a

measurement of the model and is not one.

What was actually done

Here is the whole procedure. It has no hidden steps.

FIG 1How a benchmark number is produced
Three separate human decisions sit upstream of every reported score, and the convention is to report only what comes out of the last box. Most disagreements about model quality are really disagreements about the first three boxes conducted as though they were about the last one.

So the number is a count divided by a count.

FIG 2A score is an average of ones and zeros
the published score, between zero and one
how many questions were asked
one for a question marked right, zero for one marked wrong
Writing it this way is not pedantry. It makes visible that a score is a sample average, which means every ordinary fact about sample averages applies to it, including that it would have come out differently with a different sample. The next lesson is entirely about how much differently.

That is the entire definition. There is nothing else in it. A score has no more

authority than any other average of a few hundred yes-or-no outcomes, and the

confident way scores are quoted tends to obscure that.

The questions nobody asked

Nobody actually cares about those four hundred questions. They are standing in

for something: the much larger and vaguer set of questions people might really

ask of the system. The score is useful exactly to the degree that the stand-in

is honest, and that is a judgement about how the questions were collected rather

than anything you can read off the number.

FIG 3What a hundred questions in one popular set are actually asking for
This set would be reported as a single subject score. Two thirds of it is recall or a single short step. A model that memorised well and reasoned poorly would do well here, and the name on the benchmark would not warn you.

There is a further consequence that is easy to miss. Because the set is a

stand-in, a score can be perfectly accurate about those four hundred questions

and still tell you nothing useful, if the four hundred were gathered in a way

that does not resemble the use you have in mind. A set scraped from exam papers

is honest about exam papers. Whether that transfers to the questions your

colleagues actually type is a separate claim, and it is a claim about

collection, not about the model. Nobody can settle it by running more

evaluations on the same set.

The composition is a choice, and it was made by whoever assembled the set. Shift

the proportions and the same models come out in a different order, with nobody

having done anything dishonest. This is why two benchmarks that both claim to

measure reasoning can disagree about which system reasons better.

The number belongs to three things

A lap time is not a property of a car. It is a property of a car, a track and a

driver, and quoting it without the other two is how people end up surprised.

Benchmark scores work the same way.

FIG 4Three systems across four question sets
system Asystem Bsystem C
set one889184
set two767179
set three626658
set four413347
System B wins on the first and third sets, system C on the second and fourth. Any of these rows could be the one a vendor quotes. Nobody has to lie to produce an entirely misleading impression; they only have to choose which row to print.

This is the single most useful habit to form. When you read a score, ask which

question set, under what conditions, marked how. If the answer to any of the

three is unavailable, you have been given a number without the information

needed to interpret it, and the appropriate response is not scepticism about the

model but scepticism about the claim.

The marking rule is half the instrument

The last decision is the quietest one. Somebody had to write code that looks at

what the model produced and decides whether it is right.

FIG 5The same answer, four defensible marking rules
steprulemarked rightscore out of 100what happened
116161The response must equal the reference answer exactly after trimming spaces.
227474The response must contain the reference answer somewhere in it.
337979A second model is shown both and asked whether they agree.
447777A person reads both and decides.
4 steps
Identical responses, identical questions, one model. The spread across marking rules is eighteen points, which is larger than almost every gap anyone argues about publicly. The rule that is closest to the human reading is not the strictest one, because exact match punishes answers that are right and phrased differently.

None of the four rules is wrong. Exact match is the most reproducible and the

least forgiving. Substring match is forgiving in ways that occasionally let

through an answer that contains the right string inside a wrong sentence. A

judging model is closest to a reader and brings its own biases, including a

preference for answers that look like its own writing. A human reader is the

reference everybody is approximating and is too slow to use at scale.

The point is not to find the correct rule. It is that the rule is part of the

measuring instrument, so a score quoted without it is like a length quoted

without saying whether it was measured in centimetres or inches.

What to hold on to

A benchmark score is a sample average of yes-or-no outcomes, produced jointly by

a model, a chosen set of questions and a chosen way of running and marking them.

Treating it as a property of the model drops two of the three contributors. For

the rest of this course, every time a number appears, the first question is the

same one: a sample of what, marked how, run under what conditions.

Recap

  • A score is the average of a yes or no answer over a sample of questions, which means it carries all the ordinary weaknesses of an average over a sample and none of the authority of a measured constant.
  • The number belongs to a combination of three things, the model, the question set and the conditions it was run under, and quoting it as a fact about the model alone drops two thirds of what produced it.
  • Who wrote the questions determines what the score can say, and two honest benchmarks with the same name on the box can rank the same three systems in different orders.

This is the reading half

Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.

The contents

NextThe Error Bar Nobody Prints →

The rest of this course

  1. 01The Number Describes the Questions as Much as the Modelyou are here
  2. 02How Far the Number Could Have Landed From Where It Didopening only
  3. 03Only the Questions They Disagree About Carry Any Informationopening only
  4. 04The Top of the Table Is Not Where the Best System Isopening only
  5. 05A Model That Has Seen the Exam Is Not Being Examinedopening only
  6. 06The Name on the Box Is a Claim, Not a Descriptionopening only
  7. 07Everything Between the Question and the Numberopening only
  8. 08Say the Number, the Range, the Conditions and the Limitsopening only

Read alongside