ContentsThe library

Where Training Data Comes From

What the Data Fixes, and What It Cannot

A training set sets a ceiling no later choice can lift. It also leaves several things entirely open, and confusing the two is how teams spend a year on the wrong problem.

Two questions that get confused

Ask why one model is better than another and you will be told about

architecture, parameter count and training budget. Those answers are available,

quotable and mostly wrong. The usual answer is the data, and it is rarely given

because the data is the part nobody looked at.

But "the data decides" is too broad to act on. There are things a corpus fixes

absolutely and things it has almost no bearing on, and most wasted effort comes

from treating one as the other: a team spending six months on architecture to

fix a gap that is an absence in the corpus, or a team rebuilding the corpus to

fix a latency problem that was always a serving decision.

The part that is fixed

The loss a model reaches splits cleanly into two pieces.

FIG 1Two parts of a loss
the loss actually measured
the irreducible part: genuine uncertainty in the text itself, which no model can remove
the reducible part: structure the model has not yet learned
Only the second term is for sale. More compute, a better architecture and longer training all buy reductions in R, and none of them touch H.

That floor is a property of the data, not of the model. Text that is genuinely

unpredictable stays unpredictable. This is the precise sense in which a corpus

sets a ceiling: it fixes the best achievable loss, and everything the field

spends money on is the distance between where you are and that number.

The practical version of the ceiling is less abstract. A model cannot be

competent at something no text in its corpus demonstrates. Not less competent:

not competent. There is no mechanism by which the weights acquire information

that never passed through them.

The mixture is the specification

Within what is present, competence tracks proportion.

FIG 2A corpus, by where it came from
This chart is the nearest thing the project has to a specification document, and in most projects nobody writes it down or looks at it after the first week.

A reader can predict a great deal from that chart alone. The model will be

fluent, will write serviceable code in whatever languages dominate those

repositories, will be confident and often wrong on specialist technical

questions, and will be noticeably weaker in any language that makes up a

fraction of a per cent of the web.

That is a capability forecast made before training, from the mixture, and it is

usually closer to the truth than the one made afterwards from the benchmark

scores. It is worth writing the chart down and treating it as the thing it is.

Absence leaves no trace

The hardest property of data faults is that the worst one is silent.

FIG 3Three runs, same architecture
stepruncorpusheld-out losswhat it turned out to be bad atwhat happened
1102.310General web mixture. Loss 2.31. Weak at arithmetic, which nobody noticed for two months.
2212.340The same, plus three per cent mathematics. Loss slightly worse, arithmetic much better. The loss went the wrong way and the model improved.
3322.290The same, with maths removed and web increased. Best loss of the three, worst model for the intended use.
3 steps
The loss ranked these three runs in exactly the wrong order, because it is computed over the mixture each one was trained on.

This is the trap. The loss is measured on held-out data drawn from the same

mixture, so a corpus missing something is evaluated by a measurement that is

also missing it. Nothing in a training run can report on text that was never in

it. A capability gap of this kind is found by a person asking the model to do

the thing and watching it fail, usually some months later and usually in front

of a customer.

The only defence is to write down what the model is supposed to do before

building the corpus, and then to check each item against the mixture by looking

for the text that demonstrates it. That check takes a day and is the highest

return day in the project.

What the data does not decide

Equally important, and more often got wrong in the other direction.

The corpus has essentially no say in how fast the model runs, which is set by

its size and shape. It does not set the cost of serving a request, which is the

same arithmetic regardless of what the weights contain. It does not set how much

context the model can attend to, which is an architectural limit. And it has

surprisingly little say in the format and manner of answers: whether the model

replies briefly or at length, refuses or complies, uses lists or prose, is

decided overwhelmingly by the later alignment stages working on a few tens of

thousands of examples, not by the trillion tokens underneath.

FIG 4Which stage can repair which fault
The further left a fault lies, the more it costs to fix. This is the whole argument for spending time on the corpus before spending money on the run.

The diagram is worth keeping in mind during any argument about why a model is

disappointing. The three stages fail in distinguishable ways, the cost of

repairing each differs by two orders of magnitude, and the most expensive one is

the one that is cheapest to investigate before the fact.

Recap

  • The data fixes what a model can be good at, and no amount of architecture, scale or tuning afterwards repairs an absence in it.
  • The data does not fix how fast the model runs, what it costs to serve, or the shape of its answers, all of which are decided elsewhere.
  • Whatever is missing from the corpus leaves no trace in any number the run reports, which is why absence is the most expensive fault and the last one found.

This is the reading half

Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.

The contents

NextWhere the Text Comes From →

The rest of this course

  1. 01What the Data Fixes, and What It Cannotyou are here
  2. 02Five Places, and What Each One Is Biased Towardsopening only
  3. 03Filters, In the Order That Costs Leastopening only
  4. 04Deduplication Is the Highest-Value Hour in the Projectopening only
  5. 05Setting the Proportions on Purposeopening only
  6. 06The Benchmark That Lies in the Flattering Directionopening only
  7. 07Writing a Task Two People Can Agree Onopening only
  8. 08Describing a Corpus in Numbers a Reader Can Checkopening only

Read alongside