The Curve In Every Textbook
Last timeHow Many Shapes A Model Can Take
Error on new data breaks into a part from being too rigid and a part from being too sensitive, and the two move in opposite directions as capacity grows.
There are two ways to get the wrong answer and they call for opposite remedies.
You can use a model so rigid that nothing it could possibly express resembles the
truth, in which case it will be wrong in the same direction every time. Or you can
use one so flexible that it fits the particular sample it was handed, quirks
included, in which case it will be wrong in a different direction every time you
retrain it. Capacity is the dial that trades one for the other, and the curve that
results is the oldest picture in the subject.
The two ways of being wrong
Take the first. Fit a straight line to something that genuinely curves. Every
training example is missed, the misses follow a pattern, and nothing about the
situation improves if you collect ten times as much data, because the family of
straight lines does not contain the answer. The diagnostic signature is a high
training error. The model is not failing to generalise. It never fitted in the
first place.
Now take the second. Fit a wiggly curve with twenty free quantities to twelve
noisy points. It passes through all of them, so the training error is near zero,
and it has done so by inventing structure out of the noise. Train it again on a
different twelve points from the same source and you get a visibly different
curve. The diagnostic signature is a low training error with a large gap, and
instability across retraining.
| step | family fitted | training error | error on new data | which problem |
|---|---|---|---|---|
| 1 | a constant, one free quantity | very high | very high | far too rigid, misses systematically |
| 2 | a straight line, two free quantities | high | high | still too rigid for curved truth |
| 3 | a cubic, four free quantities | low | lowest of the four | about right for this many examples |
| 4 | a polynomial of degree eleven, twelve fr | zero | very high | far too sensitive, fitted the noise |
The remedies are opposite and this is why the distinction earns its keep. Rigidity
is cured by a larger family, better inputs or a different kind of model, and is
untouched by more data or by penalties. Sensitivity is cured by more data, by
penalties, by stopping earlier or by averaging several models, and is untouched by
making the family larger.
The lesson stops here
6 more paragraphs to go
You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.
The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.
See the planThe contentsThis is the reading half
Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.
The contents