The library

How models work

Read the number a model puts next to an answer for what it actually is, measure whether it can be trusted, and repair it, well enough to decide when a system may act on its own and when it must ask

When a Model Is Sure and Wrong

Models report confidence and people believe it. This course derives what that number means, shows how to measure whether it is honest, and covers the repairs: temperature, held-out calibration, abstaining, and knowing what cannot be fixed.

8 lessons, written and corrected before you arrived. Reading them here needs no account. The first reads the whole way through; the others open and then stop, because a page nobody owns cannot tell who is reading it. Starting the course gives you your own copy, where every idea has problems standing under it and you can ask about any sentence.

Start reading

  1. 01Where the Confidence Comes FromA reported probability is the output of one fixed formula applied to a list of scores. Knowing that formula tells you exactly what the number is a probability of, which is not what people read it as.
  2. 02The Reliability Curve, in One Afternoonopening onlyWhether a model's ninety per cent means ninety per cent is a measurement, not an opinion. It takes predictions you already have, one sort, a few buckets and a plot.
  3. 03The Objective Rewards Certainty It Has Not Earnedopening onlyOverconfidence is not a defect in the model. It is what the training objective asks for, and the drift towards it continues long after accuracy has stopped improving.
  4. 04Divide Every Score by the Same Numberopening onlyOne scalar, fitted on held-out data, removes most systematic overconfidence and provably changes no answer the model gives. It is the best ratio of effect to risk in the field.
  5. 05Uncertainty More Data Would Remove, and Uncertainty It Would Notopening onlySome uncertainty is in the question and cannot be trained away. Some is in the model and would vanish with the right examples. They look identical in the output and demand opposite responses.
  6. 06Ask It Five Times and Count the Answersopening onlyDisagreement between repeated answers measures something a single probability cannot: whether the model is sure because it knows, or sure because it happened to land somewhere.
  7. 07The Threshold Comes From Your Costs, Not From a Round Numberopening onlyAnswering beats refusing exactly when the expected cost of a wrong answer falls below the cost of a referral. That comparison gives a threshold, and the threshold is where calibration pays off.
  8. 08Why a Made-Up Answer Can Carry a High Numberopening onlyA fabricated answer is fluent, and fluency is what the probability measures. Nothing in this course removes that, and it is worth knowing which repairs do nothing about it.