When a Model Is Sure and Wrong
The Objective Rewards Certainty It Has Not Earned
Last timeChecking Whether It Is Honest
Overconfidence is not a defect in the model. It is what the training objective asks for, and the drift towards it continues long after accuracy has stopped improving.
The loss never stops pushing
Models are trained to minimise the negative logarithm of the probability they
assign to the correct answer. That single choice explains most of what this
course is about.
- the loss contributed by one training example
- the probability the model gave to the answer that was actually correct
Look at the shape. At a probability of 0.9 the loss is about 0.105. At 0.99 it
is about 0.010. At 0.999 it is about 0.001. The improvement is real every time,
so there is a gradient pushing the probability up at every point short of one,
and the push never weakens into nothing.
Nothing in this asks whether the certainty is warranted. The loss on a training
example is a function of the probability given to the known answer, and the
known answer is known. Confidence that is justified and confidence that is
merely asserted earn exactly the same reduction.
The lesson stops here
4 more paragraphs to go
You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.
The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.
See the planThe contentsThis is the reading half
Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.
The contents