ContentsThe library

When a Model Is Sure and Wrong

Why a Made-Up Answer Can Carry a High Number

Last timeLetting It Say Nothing

A fabricated answer is fluent, and fluency is what the probability measures. Nothing in this course removes that, and it is worth knowing which repairs do nothing about it.

Why an invention scores high

Lesson one derived what the reported number is: a function of the scores the

model assigned to the words it chose. That derivation is the whole explanation

of this lesson, and it is worth putting plainly.

Ask for the name of the lead author of a paper that does not exist. The model

does not have a slot marked "paper exists: no". It has a partly written sentence

and a question about what comes next. Given "the lead author of that paper is",

the likely continuations are surnames, and one of them will score well above the

others because surnames are what follows that phrase. Every token in the answer

is a highly likely token, so their product is large, so the reported confidence

is high.

FIG 1How a question with no answer produces a confident one
The number is not malfunctioning. It reports how likely the words were, and a well-formed invention is made of likely words, so a high number is the right answer to the question the number asks.

So confident fabrication is not a failure of the confidence mechanism. It is the

mechanism working on a question it was never asked: whether the referent exists.

Nothing in the chain from text to score to number touches that question, and no

amount of improving the number will make it start.

What calibration cannot see

Every repair in this course improves the ordinary case. Temperature scaling

divides the scores so that the 90 per cent bucket is right 90 per cent of the

time. Asking twice detects questions the model would answer differently on a

second attempt. Both work, and both are nearly blind here.

Temperature is a single number fitted on held-out data to make the averages

line up. If fabrications are three per cent of that data, the fitted temperature

is set almost entirely by the other ninety-seven, and the fabrications are

rescaled by whatever number the majority asked for. Nothing in the fitting

notices them as a separate kind of case. Asking twice does better, because a

model sampling freely may invent a different surname each time, and disagreement

is a genuine signal. But where the invention is the single most likely

continuation, five samples produce the same one five times. The model is not

uncertain. It is consistent and wrong.

The lesson stops here

3 more paragraphs to go

You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.

The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.

See the planThe contents

This is the reading half

Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.

The contents

The rest of this course

  1. 01Where the Confidence Comes From
  2. 02The Reliability Curve, in One Afternoonopening only
  3. 03The Objective Rewards Certainty It Has Not Earnedopening only
  4. 04Divide Every Score by the Same Numberopening only
  5. 05Uncertainty More Data Would Remove, and Uncertainty It Would Notopening only
  6. 06Ask It Five Times and Count the Answersopening only
  7. 07The Threshold Comes From Your Costs, Not From a Round Numberopening only
  8. 08Why a Made-Up Answer Can Carry a High Numberyou are here

Read alongside