ContentsThe library

Why Learning Works At All

The Number That Proves Nothing

A model can score perfectly on the examples it was trained on without having learned anything, so training performance is not evidence and something else has to be measured.

Here is a model. It keeps a list of every example it was trained on along with the

right answer. Shown an input, it searches the list. If the input is there it

returns the stored answer. If it is not, it guesses.

On the training examples this model is perfect. Not good, perfect, and it will

stay perfect however many examples you add. It has also learned nothing whatsoever,

and will be exactly as useful as random guessing on anything new.

Nobody builds this model, but it establishes something important. A perfect

training score is compatible with having learned nothing. So a training score,

however good, is not evidence about anything. The number is real; the inference

people draw from it is not available.

FIG 1The memoriser, evaluated two ways
stepwhat is measuredthe memoriser scoreswhat the score shows
1accuracy on the training examples100 percentthat the list was stored correctly
2accuracy on new examples from the same schancethat nothing was learned
3accuracy after adding ten times more tra100 percentstill only that the list was stored corr
4accuracy on new examples after thatchancemore data did not help, because the meth
5the gap between the twoas large as it can bethe whole of the training score was fitt
5 steps
The third row is the one worth sitting with. More data is the usual first suggestion when a model disappoints, and here it changes nothing at all, because the failure is in what the procedure does with data rather than in how much of it there is. Telling those two situations apart is most of what this course is for.

Why the first number is contaminated

The training examples were not neutral observers. They were used to pick the

model. Every adjustment made during training was made in order to improve

performance on exactly those examples, including adjustments that fitted their

particular accidents.

Real examples carry accidents. A measurement is a little off. An unusual case got

included. Two features happen to line up in this sample and not in general. A

flexible model will fit all of it, because fitting it improves the number being

optimised, and nothing in the procedure distinguishes an accident from a pattern.

FIG 2The two numbers as a model is given more freedom
0.006.2512.5018.7525.001.05.810.515.320.0how flexible the model is
error on the training exampleserror on examples never seen
The lower curve falls forever, because more freedom always allows a closer fit to the examples in hand. The upper curve falls and then rises, because past a point the extra freedom is spent fitting things that will not recur. Reading only the lower curve, more flexibility always looks better, which is exactly the mistake the figure exists to prevent.
FIG 3The gap, named
error measured on the examples used to fit the model
error on examples that played no part in fitting it
the generalisation gap, which is how much of the training performance did not transfer
Neither number means much alone. A training error of two percent is excellent if the gap is one point and meaningless if the gap is thirty. Most of the machinery in later lessons exists either to estimate this gap before paying to find out, or to make it smaller.

It is worth being clear that this is not an exotic failure. A flexible model fitted

to a few hundred examples will usually drive its training error to near zero, and

the only question is how much of that was pattern and how much was accident. The

training number cannot answer the question because it is computed on the very data

whose accidents were fitted. Asking it is like asking a student to mark their own

paper against the answers they wrote.

The assumption nobody states

There is a condition under which any of this can work, and it is usually left

implicit. Future examples must come from the same process as past ones. Not

identical examples, but the same underlying source, so that what was true of the

sample tends to be true of the population.

Grant that and learning from examples is possible. Deny it and nothing is. If the

process generating examples changes after training, a model fitted to the old

process has no claim on the new one, and the quantity of training data is

irrelevant to the problem.

This matters in practice far more than its brief treatment in textbooks suggests,

because deployed systems live in a world that shifts under them. Prices change,

users change, the thing being measured changes, and sometimes the model itself

changes the behaviour it was predicting.

FIG 4What has to hold for a measurement to transfer
Notice what the diagram does not contain. There is no arrow from training performance to performance in use. The only path runs through a measurement on examples held out from the same process, and even that path carries a condition.

Reading the pair

Since neither number alone is informative, the practical skill is reading them

together. Low training error with a wide gap means the model fitted the specific

examples. High training error with a narrow gap means the model was too rigid to

fit anything, and the gap is narrow only because there was nothing to lose.

Those two situations call for opposite responses, and the two-number reading is

the fastest way to tell them apart.

FIG 5
The pair of numbersWhat it meansWhat to changeCommon misreadingHow often
Low training error, narrow gapit is workingnothing yetthat there is no more to getthe goal
Low training error, wide gapfitted the sampleless freedom, more datathat training harder will close itvery common
High training error, narrow gaptoo rigid, or the task is hardmore freedom, better inputsthat the task is impossiblecommon
High training error, wide gapthe setup is brokenthe measurement, not the modelthat it is a modelling problemrare

The last row is the odd one. It usually means a mismatch between the two sets, a

bug in how one of them is measured, or a held-out set too small to be stable. When

the pair makes no sense, check the measurement before changing the model.

So the first number is worthless and the second number is everything. The next

lesson is about the second number: what exactly has to be true for it to mean what

it claims, and the ordinary, easy, extremely common ways it quietly stops being

true.

Recap

  • A model that simply stored every training example would score perfectly on them, so a perfect training score is consistent with having learned nothing at all.
  • The only claim worth making is about examples the model has not seen, and nothing about the training examples can establish it.
  • Generalisation is only possible because new examples are assumed to come from the same source as old ones, and when that assumption breaks no amount of training data helps.

This is the reading half

Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.

The contents

NextKeeping Some Back →

The rest of this course

  1. 01The Number That Proves Nothingyou are here
  2. 02The Measurement And How It Breaksopening only
  3. 03The Word That Settles The Argumentopening only
  4. 04The Curve In Every Textbookopening only
  5. 05Charging For The Wrong Answersopening only
  6. 06The Lever With No Downside, Almostopening only
  7. 07Past The Point Where It Should Breakopening only
  8. 08The Honest State Of The Answeropening only

Read alongside