The Number Describes the Questions as Much as the Model
A benchmark score is a fraction of a few hundred questions somebody chose, marked by a rule somebody wrote, under conditions somebody picked. None of those three are the model.
Somebody tells you a model scores eighty four. Before anything else, it is worth
being precise about what has just been said, because the sentence sounds like a
measurement of the model and is not one.
What was actually done
Here is the whole procedure. It has no hidden steps.
So the number is a count divided by a count.
- the published score, between zero and one
- how many questions were asked
- one for a question marked right, zero for one marked wrong
That is the entire definition. There is nothing else in it. A score has no more
authority than any other average of a few hundred yes-or-no outcomes, and the
confident way scores are quoted tends to obscure that.
The questions nobody asked
Nobody actually cares about those four hundred questions. They are standing in
for something: the much larger and vaguer set of questions people might really
ask of the system. The score is useful exactly to the degree that the stand-in
is honest, and that is a judgement about how the questions were collected rather
than anything you can read off the number.
There is a further consequence that is easy to miss. Because the set is a
stand-in, a score can be perfectly accurate about those four hundred questions
and still tell you nothing useful, if the four hundred were gathered in a way
that does not resemble the use you have in mind. A set scraped from exam papers
is honest about exam papers. Whether that transfers to the questions your
colleagues actually type is a separate claim, and it is a claim about
collection, not about the model. Nobody can settle it by running more
evaluations on the same set.
The composition is a choice, and it was made by whoever assembled the set. Shift
the proportions and the same models come out in a different order, with nobody
having done anything dishonest. This is why two benchmarks that both claim to
measure reasoning can disagree about which system reasons better.
The number belongs to three things
A lap time is not a property of a car. It is a property of a car, a track and a
driver, and quoting it without the other two is how people end up surprised.
Benchmark scores work the same way.
| system A | system B | system C | |
|---|---|---|---|
| set one | 88 | 91 | 84 |
| set two | 76 | 71 | 79 |
| set three | 62 | 66 | 58 |
| set four | 41 | 33 | 47 |
This is the single most useful habit to form. When you read a score, ask which
question set, under what conditions, marked how. If the answer to any of the
three is unavailable, you have been given a number without the information
needed to interpret it, and the appropriate response is not scepticism about the
model but scepticism about the claim.
The marking rule is half the instrument
The last decision is the quietest one. Somebody had to write code that looks at
what the model produced and decides whether it is right.
| step | rule | marked right | score out of 100 | what happened |
|---|---|---|---|---|
| 1 | 1 | 61 | 61 | The response must equal the reference answer exactly after trimming spaces. |
| 2 | 2 | 74 | 74 | The response must contain the reference answer somewhere in it. |
| 3 | 3 | 79 | 79 | A second model is shown both and asked whether they agree. |
| 4 | 4 | 77 | 77 | A person reads both and decides. |
None of the four rules is wrong. Exact match is the most reproducible and the
least forgiving. Substring match is forgiving in ways that occasionally let
through an answer that contains the right string inside a wrong sentence. A
judging model is closest to a reader and brings its own biases, including a
preference for answers that look like its own writing. A human reader is the
reference everybody is approximating and is too slow to use at scale.
The point is not to find the correct rule. It is that the rule is part of the
measuring instrument, so a score quoted without it is like a length quoted
without saying whether it was measured in centimetres or inches.
What to hold on to
A benchmark score is a sample average of yes-or-no outcomes, produced jointly by
a model, a chosen set of questions and a chosen way of running and marking them.
Treating it as a property of the model drops two of the three contributors. For
the rest of this course, every time a number appears, the first question is the
same one: a sample of what, marked how, run under what conditions.
Recap
- A score is the average of a yes or no answer over a sample of questions, which means it carries all the ordinary weaknesses of an average over a sample and none of the authority of a measured constant.
- The number belongs to a combination of three things, the model, the question set and the conditions it was run under, and quoting it as a fact about the model alone drops two thirds of what produced it.
- Who wrote the questions determines what the score can say, and two honest benchmarks with the same name on the box can rank the same three systems in different orders.
This is the reading half
Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.
The contents