ContentsThe library

When a Model Is Sure and Wrong

Where the Confidence Comes From

A reported probability is the output of one fixed formula applied to a list of scores. Knowing that formula tells you exactly what the number is a probability of, which is not what people read it as.

Where the number comes from

A classifier does not produce probabilities. It produces a list of real numbers,

one per option, with no constraint on them at all: they can be negative, they can

be enormous, they do not sum to anything in particular. These are the scores.

Something then has to turn that list into numbers a person can read. The

something is one fixed formula, applied identically everywhere.

FIG 1Scores into probabilities
the score for option i, exponentiated, which makes it positive and makes a one-unit lead into a fixed multiplicative advantage
the same quantity totalled over every option offered, which is what forces the outputs to sum to exactly one
This is the whole of it. Everything the number is, and everything it is not, follows from this one line being applied to a list of unconstrained scores.

Two properties follow and no others. The outputs are positive. They sum to one.

That is what the formula guarantees. It does not guarantee, and nothing in it

attempts to guarantee, that an option given 0.9 is right nine times in ten.

FIG 2The same formula on two questions
stepquestionscoresreported probabilitiesactually right?what happened
1one the model has seen often4.1, 1.2, 0.70.92, 0.05, 0.03yesA three-unit lead becomes 0.92. The lead is large because the training data was unambiguous here.
2one it has no idea about1.4, 1.3, 1.20.37, 0.34, 0.29noScores close together, so the formula spreads the mass out. This is the honest case and it is rarer than it should be.
3one it has seen often but misleadingly3.9, 1.1, 0.90.91, 0.06, 0.03noIndistinguishable from the first row by the number alone. The formula has no access to whether the pattern it learned was true.
3 steps
Rows one and three produce almost the same number and differ in the only respect anybody cares about. Nothing downstream of the scores can tell them apart.

A probability of what

The division by the total is the part people skip, and it is where the meaning

is set. Dividing by the sum over the options offered makes the output a

conditional probability: the chance that this option is the answer, given that

one of these options is.

There is no option for none of them. If the right answer is not in the list, the

formula still returns numbers summing to one, and one of the wrong options gets

most of the mass. The model has not failed; it answered the question it was

asked, which was which of these, not whether any of these.

FIG 3The chain from text to number
The formula in the middle has no parameters and learns nothing. Whatever is wrong with the scores arrives at the reader intact, wearing the costume of a probability.

One token, not one claim

For a language model the situation is further removed from what people assume.

The model does not score answers. It scores the next token, over a vocabulary of

a hundred thousand or so. A twenty-token answer carries twenty such numbers.

Multiplying them gives the probability of that exact sequence of tokens. That is

a real quantity and it is not the probability that the answer is correct. It is

the probability of one wording. The same claim phrased differently is a different

sequence with a different and usually much smaller number, and a longer answer is

mechanically less probable than a shorter one, which is why averaging per token

rather than multiplying is the usual repair and why the result is still not a

chance of being right.

FIG 4How a score lead becomes a probability
0.000.300.600.901.200.01.53.04.56.0how far ahead the leading score is
two optionsten options
A lead of three units reports about 0.95 with two options and about 0.75 with ten. The same internal judgement produces different numbers depending only on how many options were offered, which is a property of the formula rather than of the evidence.

Ordering is not magnitude

The salvage is that the ordering is usually good. Training rewarded putting the

right answer above the wrong ones, so the scores rank well, and anything that

only needs the ranking, which includes picking the top answer and returning a

shortlist, works.

Nothing in training constrained the size of the gaps. The loss is minimised by

pushing the right score up without limit, and a model that is right ninety per

cent of the time can report 0.99 throughout without paying for it anywhere.

So the practical rule is to use the number as an ordering and never as a rate

until it has been measured and repaired. Using it as a rate is what produces the

system that routes anything above 0.9 straight through to a customer, on the

belief that one in ten will be wrong, when in fact one in three is.

Recap

  • The number is produced by a formula that forces a list of scores to sum to one. It is a probability over the options offered, not over the truth.
  • For a language model the number is attached to one token, and a sentence is a product of many of them, which is a different quantity entirely.
  • The ordering the scores produce is usually reliable; the magnitudes usually are not, and almost every misuse is reading a magnitude as a rate.

This is the reading half

Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.

The contents

NextChecking Whether It Is Honest →

The rest of this course

  1. 01Where the Confidence Comes Fromyou are here
  2. 02The Reliability Curve, in One Afternoonopening only
  3. 03The Objective Rewards Certainty It Has Not Earnedopening only
  4. 04Divide Every Score by the Same Numberopening only
  5. 05Uncertainty More Data Would Remove, and Uncertainty It Would Notopening only
  6. 06Ask It Five Times and Count the Answersopening only
  7. 07The Threshold Comes From Your Costs, Not From a Round Numberopening only
  8. 08Why a Made-Up Answer Can Carry a High Numberopening only

Read alongside