ContentsThe library

When a Model Is Sure and Wrong

Ask It Five Times and Count the Answers

Last timeTwo Different Kinds of Not Knowing

Disagreement between repeated answers measures something a single probability cannot: whether the model is sure because it knows, or sure because it happened to land somewhere.

Asking repeatedly

Decoding is random whenever the temperature is above zero. The same question,

asked five times, produces five runs through the same model with different

random choices at each step, and therefore sometimes five different answers.

The single-pass workflow throws that variation away. It asks once, takes what

came back, and reports the probability attached to it. But the variation is

information: it says how much of the answer was determined by what the model

knows and how much by which way a coin fell.

FIG 1One question, five askings
stepaskinganswerprobability it reportedwhat happened
1119470.91Confident, and so far there is nothing to compare it against.
2219470.88Same answer. Two for two.
3319470.93Three.
4419510.84A different answer, reported almost as confidently. The single-pass view would never have seen this.
5519470.9Four of five agree. The agreement figure is 0.80, against a mean reported probability of 0.89.
5 steps
The reported probabilities barely move across the five runs, which is the point: they are not measuring what the disagreement measures. On a question the model actually knows, all five would say the same thing.

Agreement as a number

Collapse the samples into one figure by counting.

FIG 2The agreement fraction
how many times the question was asked
how many of those landed on the most common answer
the agreement fraction, between one over k and one
Counting rather than computing. Nothing in this came out of the model's own scoring, which is exactly why it carries information the reported probability does not.

Two details make it work in practice. Answers have to be compared after

normalising, since 1947 and the year 1947 are the same answer and a string

comparison says otherwise; for open-ended answers this means comparing the claim

rather than the wording, which is usually a second model's job. And k has to be

small enough to afford and large enough to mean something: three is the

practical floor, five is the common choice, and past about ten the figure stops

moving.

The lesson stops here

4 more paragraphs to go

You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.

The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.

See the planThe contents

This is the reading half

Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.

The contents

The rest of this course

  1. 01Where the Confidence Comes From
  2. 02The Reliability Curve, in One Afternoonopening only
  3. 03The Objective Rewards Certainty It Has Not Earnedopening only
  4. 04Divide Every Score by the Same Numberopening only
  5. 05Uncertainty More Data Would Remove, and Uncertainty It Would Notopening only
  6. 06Ask It Five Times and Count the Answersyou are here
  7. 07The Threshold Comes From Your Costs, Not From a Round Numberopening only
  8. 08Why a Made-Up Answer Can Carry a High Numberopening only

Read alongside