Where Training Data Comes From
What the Data Fixes, and What It Cannot
A training set sets a ceiling no later choice can lift. It also leaves several things entirely open, and confusing the two is how teams spend a year on the wrong problem.
Two questions that get confused
Ask why one model is better than another and you will be told about
architecture, parameter count and training budget. Those answers are available,
quotable and mostly wrong. The usual answer is the data, and it is rarely given
because the data is the part nobody looked at.
But "the data decides" is too broad to act on. There are things a corpus fixes
absolutely and things it has almost no bearing on, and most wasted effort comes
from treating one as the other: a team spending six months on architecture to
fix a gap that is an absence in the corpus, or a team rebuilding the corpus to
fix a latency problem that was always a serving decision.
The part that is fixed
The loss a model reaches splits cleanly into two pieces.
- the loss actually measured
- the irreducible part: genuine uncertainty in the text itself, which no model can remove
- the reducible part: structure the model has not yet learned
That floor is a property of the data, not of the model. Text that is genuinely
unpredictable stays unpredictable. This is the precise sense in which a corpus
sets a ceiling: it fixes the best achievable loss, and everything the field
spends money on is the distance between where you are and that number.
The practical version of the ceiling is less abstract. A model cannot be
competent at something no text in its corpus demonstrates. Not less competent:
not competent. There is no mechanism by which the weights acquire information
that never passed through them.
The mixture is the specification
Within what is present, competence tracks proportion.
A reader can predict a great deal from that chart alone. The model will be
fluent, will write serviceable code in whatever languages dominate those
repositories, will be confident and often wrong on specialist technical
questions, and will be noticeably weaker in any language that makes up a
fraction of a per cent of the web.
That is a capability forecast made before training, from the mixture, and it is
usually closer to the truth than the one made afterwards from the benchmark
scores. It is worth writing the chart down and treating it as the thing it is.
Absence leaves no trace
The hardest property of data faults is that the worst one is silent.
| step | run | corpus | held-out loss | what it turned out to be bad at | what happened |
|---|---|---|---|---|---|
| 1 | 1 | 0 | 2.31 | 0 | General web mixture. Loss 2.31. Weak at arithmetic, which nobody noticed for two months. |
| 2 | 2 | 1 | 2.34 | 0 | The same, plus three per cent mathematics. Loss slightly worse, arithmetic much better. The loss went the wrong way and the model improved. |
| 3 | 3 | 2 | 2.29 | 0 | The same, with maths removed and web increased. Best loss of the three, worst model for the intended use. |
This is the trap. The loss is measured on held-out data drawn from the same
mixture, so a corpus missing something is evaluated by a measurement that is
also missing it. Nothing in a training run can report on text that was never in
it. A capability gap of this kind is found by a person asking the model to do
the thing and watching it fail, usually some months later and usually in front
of a customer.
The only defence is to write down what the model is supposed to do before
building the corpus, and then to check each item against the mixture by looking
for the text that demonstrates it. That check takes a day and is the highest
return day in the project.
What the data does not decide
Equally important, and more often got wrong in the other direction.
The corpus has essentially no say in how fast the model runs, which is set by
its size and shape. It does not set the cost of serving a request, which is the
same arithmetic regardless of what the weights contain. It does not set how much
context the model can attend to, which is an architectural limit. And it has
surprisingly little say in the format and manner of answers: whether the model
replies briefly or at length, refuses or complies, uses lists or prose, is
decided overwhelmingly by the later alignment stages working on a few tens of
thousands of examples, not by the trillion tokens underneath.
The diagram is worth keeping in mind during any argument about why a model is
disappointing. The three stages fail in distinguishable ways, the cost of
repairing each differs by two orders of magnitude, and the most expensive one is
the one that is cheapest to investigate before the fact.
Recap
- The data fixes what a model can be good at, and no amount of architecture, scale or tuning afterwards repairs an absence in it.
- The data does not fix how fast the model runs, what it costs to serve, or the shape of its answers, all of which are decided elsewhere.
- Whatever is missing from the corpus leaves no trace in any number the run reports, which is why absence is the most expensive fault and the last one found.
This is the reading half
Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.
The contents