When a Model Is Sure and Wrong
Where the Confidence Comes From
A reported probability is the output of one fixed formula applied to a list of scores. Knowing that formula tells you exactly what the number is a probability of, which is not what people read it as.
Where the number comes from
A classifier does not produce probabilities. It produces a list of real numbers,
one per option, with no constraint on them at all: they can be negative, they can
be enormous, they do not sum to anything in particular. These are the scores.
Something then has to turn that list into numbers a person can read. The
something is one fixed formula, applied identically everywhere.
- the score for option i, exponentiated, which makes it positive and makes a one-unit lead into a fixed multiplicative advantage
- the same quantity totalled over every option offered, which is what forces the outputs to sum to exactly one
Two properties follow and no others. The outputs are positive. They sum to one.
That is what the formula guarantees. It does not guarantee, and nothing in it
attempts to guarantee, that an option given 0.9 is right nine times in ten.
| step | question | scores | reported probabilities | actually right? | what happened |
|---|---|---|---|---|---|
| 1 | one the model has seen often | 4.1, 1.2, 0.7 | 0.92, 0.05, 0.03 | yes | A three-unit lead becomes 0.92. The lead is large because the training data was unambiguous here. |
| 2 | one it has no idea about | 1.4, 1.3, 1.2 | 0.37, 0.34, 0.29 | no | Scores close together, so the formula spreads the mass out. This is the honest case and it is rarer than it should be. |
| 3 | one it has seen often but misleadingly | 3.9, 1.1, 0.9 | 0.91, 0.06, 0.03 | no | Indistinguishable from the first row by the number alone. The formula has no access to whether the pattern it learned was true. |
A probability of what
The division by the total is the part people skip, and it is where the meaning
is set. Dividing by the sum over the options offered makes the output a
conditional probability: the chance that this option is the answer, given that
one of these options is.
There is no option for none of them. If the right answer is not in the list, the
formula still returns numbers summing to one, and one of the wrong options gets
most of the mass. The model has not failed; it answered the question it was
asked, which was which of these, not whether any of these.
One token, not one claim
For a language model the situation is further removed from what people assume.
The model does not score answers. It scores the next token, over a vocabulary of
a hundred thousand or so. A twenty-token answer carries twenty such numbers.
Multiplying them gives the probability of that exact sequence of tokens. That is
a real quantity and it is not the probability that the answer is correct. It is
the probability of one wording. The same claim phrased differently is a different
sequence with a different and usually much smaller number, and a longer answer is
mechanically less probable than a shorter one, which is why averaging per token
rather than multiplying is the usual repair and why the result is still not a
chance of being right.
Ordering is not magnitude
The salvage is that the ordering is usually good. Training rewarded putting the
right answer above the wrong ones, so the scores rank well, and anything that
only needs the ranking, which includes picking the top answer and returning a
shortlist, works.
Nothing in training constrained the size of the gaps. The loss is minimised by
pushing the right score up without limit, and a model that is right ninety per
cent of the time can report 0.99 throughout without paying for it anywhere.
So the practical rule is to use the number as an ordering and never as a rate
until it has been measured and repaired. Using it as a rate is what produces the
system that routes anything above 0.9 straight through to a customer, on the
belief that one in ten will be wrong, when in fact one in three is.
Recap
- The number is produced by a formula that forces a list of scores to sum to one. It is a probability over the options offered, not over the truth.
- For a language model the number is attached to one token, and a sentence is a product of many of them, which is a different quantity entirely.
- The ordering the scores produce is usually reliable; the magnitudes usually are not, and almost every misuse is reading a magnitude as a rate.
This is the reading half
Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.
The contents