The Number That Proves Nothing
A model can score perfectly on the examples it was trained on without having learned anything, so training performance is not evidence and something else has to be measured.
Here is a model. It keeps a list of every example it was trained on along with the
right answer. Shown an input, it searches the list. If the input is there it
returns the stored answer. If it is not, it guesses.
On the training examples this model is perfect. Not good, perfect, and it will
stay perfect however many examples you add. It has also learned nothing whatsoever,
and will be exactly as useful as random guessing on anything new.
Nobody builds this model, but it establishes something important. A perfect
training score is compatible with having learned nothing. So a training score,
however good, is not evidence about anything. The number is real; the inference
people draw from it is not available.
| step | what is measured | the memoriser scores | what the score shows |
|---|---|---|---|
| 1 | accuracy on the training examples | 100 percent | that the list was stored correctly |
| 2 | accuracy on new examples from the same s | chance | that nothing was learned |
| 3 | accuracy after adding ten times more tra | 100 percent | still only that the list was stored corr |
| 4 | accuracy on new examples after that | chance | more data did not help, because the meth |
| 5 | the gap between the two | as large as it can be | the whole of the training score was fitt |
Why the first number is contaminated
The training examples were not neutral observers. They were used to pick the
model. Every adjustment made during training was made in order to improve
performance on exactly those examples, including adjustments that fitted their
particular accidents.
Real examples carry accidents. A measurement is a little off. An unusual case got
included. Two features happen to line up in this sample and not in general. A
flexible model will fit all of it, because fitting it improves the number being
optimised, and nothing in the procedure distinguishes an accident from a pattern.
- error measured on the examples used to fit the model
- error on examples that played no part in fitting it
- the generalisation gap, which is how much of the training performance did not transfer
It is worth being clear that this is not an exotic failure. A flexible model fitted
to a few hundred examples will usually drive its training error to near zero, and
the only question is how much of that was pattern and how much was accident. The
training number cannot answer the question because it is computed on the very data
whose accidents were fitted. Asking it is like asking a student to mark their own
paper against the answers they wrote.
The assumption nobody states
There is a condition under which any of this can work, and it is usually left
implicit. Future examples must come from the same process as past ones. Not
identical examples, but the same underlying source, so that what was true of the
sample tends to be true of the population.
Grant that and learning from examples is possible. Deny it and nothing is. If the
process generating examples changes after training, a model fitted to the old
process has no claim on the new one, and the quantity of training data is
irrelevant to the problem.
This matters in practice far more than its brief treatment in textbooks suggests,
because deployed systems live in a world that shifts under them. Prices change,
users change, the thing being measured changes, and sometimes the model itself
changes the behaviour it was predicting.
Reading the pair
Since neither number alone is informative, the practical skill is reading them
together. Low training error with a wide gap means the model fitted the specific
examples. High training error with a narrow gap means the model was too rigid to
fit anything, and the gap is narrow only because there was nothing to lose.
Those two situations call for opposite responses, and the two-number reading is
the fastest way to tell them apart.
| The pair of numbers | What it means | What to change | Common misreading | How often |
|---|---|---|---|---|
| Low training error, narrow gap | it is working | nothing yet | that there is no more to get | the goal |
| Low training error, wide gap | fitted the sample | less freedom, more data | that training harder will close it | very common |
| High training error, narrow gap | too rigid, or the task is hard | more freedom, better inputs | that the task is impossible | common |
| High training error, wide gap | the setup is broken | the measurement, not the model | that it is a modelling problem | rare |
The last row is the odd one. It usually means a mismatch between the two sets, a
bug in how one of them is measured, or a held-out set too small to be stable. When
the pair makes no sense, check the measurement before changing the model.
So the first number is worthless and the second number is everything. The next
lesson is about the second number: what exactly has to be true for it to mean what
it claims, and the ordinary, easy, extremely common ways it quietly stops being
true.
Recap
- A model that simply stored every training example would score perfectly on them, so a perfect training score is consistent with having learned nothing at all.
- The only claim worth making is about examples the model has not seen, and nothing about the training examples can establish it.
- Generalisation is only possible because new examples are assumed to come from the same source as old ones, and when that assumption breaks no amount of training data helps.
This is the reading half
Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.
The contents