The Measurement And How It Breaks
Last timeMemorising the Answer Sheet
Holding examples back measures what you want only if they truly had no influence, which fails through leakage, through reuse, and through splitting on the wrong unit.
The fix for the previous lesson seems obvious. Set some examples aside before
training, never let the model near them, and measure on those at the end. The
number you get estimates performance on the underlying process, and it is the only
honest number in the whole exercise.
That is correct, and the condition attached to it is stricter than it first
appears. The held-out examples must have influenced nothing. Not the weights, not
the choice of model, not the decision about when to stop, not which features were
built. Any influence at all is a form of fitting, and fitting to the measuring
instrument is how the measurement stops working.
The ways influence arrives
Nobody trains on the test set deliberately. Influence arrives through the side
door, and the common routes are worth knowing by name because none of them
announce themselves.
Preparing the data before splitting is the most frequent. Scaling every column by
the overall average, filling missing values from the whole file, choosing which
features to keep by looking at all the rows: each of these lets information from
the held-out examples into the training process. The held-out score then comes out
slightly too good, and the margin is invisible.
Duplicates are the second. Real datasets contain the same record twice, or two
records that are near-copies, and a random split cheerfully puts one on each side.
The model recalls rather than generalises and the measurement cannot tell.
| step | step | what was done | what it cost |
|---|---|---|---|
| 1 | prepare | columns scaled using the average over ev | the training process now knows something |
| 2 | split | rows divided at random into training and | the split looks clean and is not |
| 3 | train | the model is fitted on the training rows | nothing visibly wrong |
| 4 | measure | the held-out score comes out a point or | the error is small and entirely invisibl |
| 5 | deploy | performance is worse than the report pro | the gap is attributed to the world chang |
Each look costs something
Suppose you train fifty variants and pick the one with the best held-out score.
That score is now optimistic, because the best of fifty noisy measurements is
higher than the true value of the best variant. You have not trained on the
held-out set, but you have selected on it, and selection is fitting.
The effect is small for a handful of comparisons and large for the hundreds that
accumulate in a real project. It is also the honest explanation for a familiar
pattern where a long sequence of improvements on a benchmark fails to reproduce
anywhere else.
The lesson stops here
3 more paragraphs to go
You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.
The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.
See the planThe contentsThis is the reading half
Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.
The contents