The Lever With No Downside, Almost
Last timeMaking Some Answers Harder To Reach
More examples reduce sensitivity without adding rigidity, but the gain follows a curve that flattens, and only one of the two failure modes responds at all.
Of all the levers in this course, more examples is the one with the best reputation,
and the reputation is mostly deserved. It reduces sensitivity without introducing
rigidity, which no penalty can claim, and it does so without any quantity needing to
be tuned. What it is not is universal. It addresses exactly one of the two ways of
being wrong, and the gain it delivers shrinks in a way that can be measured in
advance.
Which term it touches
The sensitivity term measures how far a fit on one sample wanders from the average
over all samples of that size. Every extra example constrains the fit further, so
the wandering shrinks, and in the simplest cases it shrinks in proportion to one over
the number of examples. That is the entire mechanism, and it explains both what more
data fixes and what it cannot.
The rigidity term is untouched. If the family cannot express the pattern, a million
examples establish the same inability as a thousand did, only with more confidence. A
straight line fitted to a curve through a million points is still a straight line.
This is why the diagnostic from the tradeoff lesson has to come first: if the training
error is already high, collecting data is the wrong project.
| Situation | Dominant term | Does more data help | What helps instead | Cheap check |
|---|---|---|---|---|
| Training high, held-out similar | the systematic miss | barely | a larger family, better inputs | fit on a quarter and a half |
| Training near zero, held-out much higher | the wandering | yes, most of all | a charge on size | fit on a quarter and a half |
| Curve flat over three doublings | the noise floor | no | better labels, better inputs | the flat curve is the check |
| Many rows, very few sources | the wandering, hidden | only new sources help | collecting more widely | split by source |
The last row catches teams out, because the row count looks enormous and the thing
behaves like a small dataset. The cheap check is the same everywhere: fit on a
quarter, a half and all of the data you already hold, and look at the three numbers
before commissioning any collection at all.
The lesson stops here
4 more paragraphs to go
You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.
The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.
See the planThe contentsThis is the reading half
Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.
The contents