ContentsThe library

Why Learning Works At All

The Measurement And How It Breaks

Last timeMemorising the Answer Sheet

Holding examples back measures what you want only if they truly had no influence, which fails through leakage, through reuse, and through splitting on the wrong unit.

The fix for the previous lesson seems obvious. Set some examples aside before

training, never let the model near them, and measure on those at the end. The

number you get estimates performance on the underlying process, and it is the only

honest number in the whole exercise.

That is correct, and the condition attached to it is stricter than it first

appears. The held-out examples must have influenced nothing. Not the weights, not

the choice of model, not the decision about when to stop, not which features were

built. Any influence at all is a form of fitting, and fitting to the measuring

instrument is how the measurement stops working.

FIG 1Three piles, three jobs
Two piles are enough when only a couple of decisions are made. Three are needed when the project will run for months and hundreds of choices will be made along the way. The question to ask is how many times the set has been consulted, because that number is roughly how much of the final score is luck.

The ways influence arrives

Nobody trains on the test set deliberately. Influence arrives through the side

door, and the common routes are worth knowing by name because none of them

announce themselves.

Preparing the data before splitting is the most frequent. Scaling every column by

the overall average, filling missing values from the whole file, choosing which

features to keep by looking at all the rows: each of these lets information from

the held-out examples into the training process. The held-out score then comes out

slightly too good, and the margin is invisible.

Duplicates are the second. Real datasets contain the same record twice, or two

records that are near-copies, and a random split cheerfully puts one on each side.

The model recalls rather than generalises and the measurement cannot tell.

FIG 2One leak, followed to its consequence
stepstepwhat was donewhat it cost
1preparecolumns scaled using the average over evthe training process now knows something
2splitrows divided at random into training andthe split looks clean and is not
3trainthe model is fitted on the training rowsnothing visibly wrong
4measurethe held-out score comes out a point or the error is small and entirely invisibl
5deployperformance is worse than the report prothe gap is attributed to the world chang
5 steps
Note the fifth row. The leak produces a modest overstatement, not a dramatic one, which is precisely why it survives review. A result that is too good gets questioned. A result that is slightly too good gets shipped, and the shortfall afterwards is blamed on something else.

Each look costs something

Suppose you train fifty variants and pick the one with the best held-out score.

That score is now optimistic, because the best of fifty noisy measurements is

higher than the true value of the best variant. You have not trained on the

held-out set, but you have selected on it, and selection is fitting.

The effect is small for a handful of comparisons and large for the hundreds that

accumulate in a real project. It is also the honest explanation for a familiar

pattern where a long sequence of improvements on a benchmark fails to reproduce

anywhere else.

The lesson stops here

3 more paragraphs to go

You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.

The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.

See the planThe contents

This is the reading half

Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.

The contents

The rest of this course

  1. 01The Number That Proves Nothing
  2. 02The Measurement And How It Breaksyou are here
  3. 03The Word That Settles The Argumentopening only
  4. 04The Curve In Every Textbookopening only
  5. 05Charging For The Wrong Answersopening only
  6. 06The Lever With No Downside, Almostopening only
  7. 07Past The Point Where It Should Breakopening only
  8. 08The Honest State Of The Answeropening only

Read alongside