Where Training Data Comes From
Describing a Corpus in Numbers a Reader Can Check
Last timeThe Part People Have to Write
A finished corpus needs a short, honest description: its size, its composition, how it was filtered, who labelled what, and which claims about the model those numbers do not support.
The numbers worth reporting
Corpora are almost always described by a single number: one trillion tokens, two
trillion, fifteen. That number is the least informative thing about the corpus.
It constrains the size of model the data can support and nothing else. Two
corpora of the same size can differ so much in composition that models trained
on them are not comparable in any respect.
A description worth reading has roughly a dozen numbers in it.
| documents, millions | tokens, trillions | duplicates removed, per | benchmark overlap, per c | |
|---|---|---|---|---|
| web crawl | 412.00 | 1.20 | 46.00 | 0.09 |
| books | 31.00 | 0.21 | 8.00 | 0.02 |
| code | 18.00 | 0.34 | 3.00 | 0.00 |
| reference | 9.00 | 0.11 | 1.00 | 0.01 |
The figures that carry information are these. Composition, as the share of
tokens from each source, which is the single most predictive number about model
behaviour. Filter survival, the fraction of raw bytes that reached the final
corpus, which says how aggressive the pipeline was. Duplication removed, by
source, which says whether the deduplication actually ran. Benchmark overlap, by
benchmark, which decides whether the evaluations mean anything. Document length
distribution, not the mean, because the mean hides a tail of ten-word pages.
Language shares. Date range, because a corpus with no documents after a certain
year cannot know anything that happened later. And the labelled portion
separately: how many items, labelled by whom, under what instructions, at what
agreement.
The description is a document
Each of those numbers is nearly free while the pipeline is running, because the
stage that computes it is already touching every document. Each is expensive or
impossible afterwards. Nobody re-reads two trillion tokens to find out what the
length distribution was, so the honest answer six months later is that it is not
known.
So the description is assembled as the corpus is built: each stage writes its own
counts as it finishes, and the document is the collection of them rather than a
study commissioned at the end. It is worth fixing the set of questions in advance
so that stages know what to record. The useful set is: what is in it, where each
part came from, under what licence or permission, what was removed and by which
rule, what was labelled and by whom, what overlap with evaluations was measured,
what the intended use is, and what uses the authors consider unsupported.
The lesson stops here
3 more paragraphs to go
You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.
The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.
See the planThe contentsThis is the reading half
Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.
The contents