When a Model Is Sure and Wrong
The Reliability Curve, in One Afternoon
Last timeWhat That Number Actually Is
Whether a model's ninety per cent means ninety per cent is a measurement, not an opinion. It takes predictions you already have, one sort, a few buckets and a plot.
The measurement
A model claims a rate. The question is whether that rate is observed. Both sides
of the comparison are available in any system that has been running: the claimed
number is stored with each prediction, and the outcome is whatever actually
happened.
The obstacle is that accuracy is not defined for a single prediction. One answer
is right or wrong; it is not seventy per cent right. So the claimed rate can only
be checked against groups. Take every prediction where the model said roughly
0.7 and ask what fraction of those were right. That is the whole method.
| predictions | mean claimed | observed accuracy | gap, points | |
|---|---|---|---|---|
| 0.50 to 0.60 | 412.00 | 0.56 | 0.54 | 2.00 |
| 0.60 to 0.70 | 638.00 | 0.67 | 0.61 | 6.00 |
| 0.70 to 0.85 | 1104.00 | 0.77 | 0.66 | 11.00 |
| 0.85 to 1.00 | 1846.00 | 0.93 | 0.78 | 15.00 |
Three details decide whether the measurement is worth anything. The outcomes
have to be real outcomes rather than a second model's opinion, or the curve
measures agreement between two systems and calls it truth. The predictions have
to come from the same traffic the system actually serves, because a curve built
on a tidy test set says nothing about a production distribution that drifted six
months ago. And the predictions have to be ones the model was not fitted on, for
the same reason a model's training accuracy is not its accuracy.
Plotted, that table is the standard picture.
The buckets are a choice
Because accuracy only exists in groups, the curve cannot be drawn without
choosing groups, and the choice changes the picture.
Too few buckets and the fault averages away. Two buckets over a model that is
honest below 0.6 and badly overconfident above 0.9 will show a modest gap
everywhere and nothing alarming anywhere. Too many and each bucket holds a
handful of predictions, so its observed accuracy is dominated by chance and the
curve becomes a scatter nobody can read.
Two practical rules cover most cases. Make the buckets equal in count rather than
equal in width, because predictions pile up near the top and equal-width buckets
leave the interesting region with almost nothing in it. And put a floor under the
bucket size, around a hundred predictions, below which the observed accuracy is
not a measurement. A model with four thousand predictions therefore supports
maybe ten buckets, and that is the honest resolution available.
The lesson stops here
4 more paragraphs to go
You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.
The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.
See the planThe contentsThis is the reading half
Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.
The contents