ContentsThe library

When a Model Is Sure and Wrong

The Reliability Curve, in One Afternoon

Last timeWhat That Number Actually Is

Whether a model's ninety per cent means ninety per cent is a measurement, not an opinion. It takes predictions you already have, one sort, a few buckets and a plot.

The measurement

A model claims a rate. The question is whether that rate is observed. Both sides

of the comparison are available in any system that has been running: the claimed

number is stored with each prediction, and the outcome is whatever actually

happened.

The obstacle is that accuracy is not defined for a single prediction. One answer

is right or wrong; it is not seventy per cent right. So the claimed rate can only

be checked against groups. Take every prediction where the model said roughly

0.7 and ask what fraction of those were right. That is the whole method.

FIG 1Four thousand predictions, bucketed
predictionsmean claimedobserved accuracygap, points
0.50 to 0.60412.000.560.542.00
0.60 to 0.70638.000.670.616.00
0.70 to 0.851104.000.770.6611.00
0.85 to 1.001846.000.930.7815.00
The fault is not uniform. At low confidence the model is nearly honest; the overstatement grows with the claim, and the worst bucket is also the largest, which is where the system sends things through unchecked.

Three details decide whether the measurement is worth anything. The outcomes

have to be real outcomes rather than a second model's opinion, or the curve

measures agreement between two systems and calls it truth. The predictions have

to come from the same traffic the system actually serves, because a curve built

on a tidy test set says nothing about a production distribution that drifted six

months ago. And the predictions have to be ones the model was not fitted on, for

the same reason a model's training accuracy is not its accuracy.

Plotted, that table is the standard picture.

FIG 2Claimed against observed
0.00.51.00.00.51.0the big bucket: claims 0.93, delivers 0.78what the model claimedhow often it was right
perfectly honestthe model abovean underconfident model
The diagonal is the only line where the claim matches the outcome. Below it the model overstates, which is the common direction; above it the model understates, which is rarer and costs you throughput rather than trust.

The buckets are a choice

Because accuracy only exists in groups, the curve cannot be drawn without

choosing groups, and the choice changes the picture.

Too few buckets and the fault averages away. Two buckets over a model that is

honest below 0.6 and badly overconfident above 0.9 will show a modest gap

everywhere and nothing alarming anywhere. Too many and each bucket holds a

handful of predictions, so its observed accuracy is dominated by chance and the

curve becomes a scatter nobody can read.

Two practical rules cover most cases. Make the buckets equal in count rather than

equal in width, because predictions pile up near the top and equal-width buckets

leave the interesting region with almost nothing in it. And put a floor under the

bucket size, around a hundred predictions, below which the observed accuracy is

not a measurement. A model with four thousand predictions therefore supports

maybe ten buckets, and that is the honest resolution available.

The lesson stops here

4 more paragraphs to go

You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.

The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.

See the planThe contents

This is the reading half

Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.

The contents

The rest of this course

  1. 01Where the Confidence Comes From
  2. 02The Reliability Curve, in One Afternoonyou are here
  3. 03The Objective Rewards Certainty It Has Not Earnedopening only
  4. 04Divide Every Score by the Same Numberopening only
  5. 05Uncertainty More Data Would Remove, and Uncertainty It Would Notopening only
  6. 06Ask It Five Times and Count the Answersopening only
  7. 07The Threshold Comes From Your Costs, Not From a Round Numberopening only
  8. 08Why a Made-Up Answer Can Carry a High Numberopening only

Read alongside