When a Model Is Sure and Wrong
Divide Every Score by the Same Number
Last timeWhy Training Makes It Worse
One scalar, fitted on held-out data, removes most systematic overconfidence and provably changes no answer the model gives. It is the best ratio of effect to risk in the field.
Dividing the scores
The fault described in the last lesson is a systematic one: the scores are
spread too far apart, everywhere, in the same direction. A systematic fault
invites a systematic correction, and the simplest available one is to divide
every score by the same number before the formula is applied.
- the score for option i, exactly as the model produced it
- the temperature, one positive number shared by every option and every input
- the corrected probability actually reported to the reader
A temperature above one shrinks the gaps between scores, so the exponentials
come out closer together and the distribution flattens. Below one it stretches
the gaps and sharpens the distribution. At exactly one, nothing happens.
Since overconfidence is the usual direction of the fault, the fitted value is
usually above one, commonly between 1.2 and 3.
The lesson stops here
4 more paragraphs to go
You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.
The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.
See the planThe contentsThis is the reading half
Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.
The contents