ContentsThe library

Why Learning Works At All

The Curve In Every Textbook

Last timeHow Many Shapes A Model Can Take

Error on new data breaks into a part from being too rigid and a part from being too sensitive, and the two move in opposite directions as capacity grows.

There are two ways to get the wrong answer and they call for opposite remedies.

You can use a model so rigid that nothing it could possibly express resembles the

truth, in which case it will be wrong in the same direction every time. Or you can

use one so flexible that it fits the particular sample it was handed, quirks

included, in which case it will be wrong in a different direction every time you

retrain it. Capacity is the dial that trades one for the other, and the curve that

results is the oldest picture in the subject.

The two ways of being wrong

Take the first. Fit a straight line to something that genuinely curves. Every

training example is missed, the misses follow a pattern, and nothing about the

situation improves if you collect ten times as much data, because the family of

straight lines does not contain the answer. The diagnostic signature is a high

training error. The model is not failing to generalise. It never fitted in the

first place.

Now take the second. Fit a wiggly curve with twenty free quantities to twelve

noisy points. It passes through all of them, so the training error is near zero,

and it has done so by inventing structure out of the noise. Train it again on a

different twelve points from the same source and you get a visibly different

curve. The diagnostic signature is a low training error with a large gap, and

instability across retraining.

FIG 1One dataset, three families
stepfamily fittedtraining errorerror on new datawhich problem
1a constant, one free quantityvery highvery highfar too rigid, misses systematically
2a straight line, two free quantitieshighhighstill too rigid for curved truth
3a cubic, four free quantitieslowlowest of the fourabout right for this many examples
4a polynomial of degree eleven, twelve frzerovery highfar too sensitive, fitted the noise
4 steps
The two outer rows are the two failures, and reading the second and third columns together tells you which you have. High on both means rigidity. Near zero on the first and high on the second means sensitivity. The middle rows are the interesting region, where the two effects are comparable in size and the choice is genuinely a trade.

The remedies are opposite and this is why the distinction earns its keep. Rigidity

is cured by a larger family, better inputs or a different kind of model, and is

untouched by more data or by penalties. Sensitivity is cured by more data, by

penalties, by stopping earlier or by averaging several models, and is untouched by

making the family larger.

The lesson stops here

6 more paragraphs to go

You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.

The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.

See the planThe contents

This is the reading half

Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.

The contents

The rest of this course

  1. 01The Number That Proves Nothing
  2. 02The Measurement And How It Breaksopening only
  3. 03The Word That Settles The Argumentopening only
  4. 04The Curve In Every Textbookyou are here
  5. 05Charging For The Wrong Answersopening only
  6. 06The Lever With No Downside, Almostopening only
  7. 07Past The Point Where It Should Breakopening only
  8. 08The Honest State Of The Answeropening only

Read alongside