ContentsThe library

Where Training Data Comes From

Describing a Corpus in Numbers a Reader Can Check

Last timeThe Part People Have to Write

A finished corpus needs a short, honest description: its size, its composition, how it was filtered, who labelled what, and which claims about the model those numbers do not support.

The numbers worth reporting

Corpora are almost always described by a single number: one trillion tokens, two

trillion, fifteen. That number is the least informative thing about the corpus.

It constrains the size of model the data can support and nothing else. Two

corpora of the same size can differ so much in composition that models trained

on them are not comparable in any respect.

A description worth reading has roughly a dozen numbers in it.

FIG 1A corpus, described
documents, millionstokens, trillionsduplicates removed, per benchmark overlap, per c
web crawl412.001.2046.000.09
books31.000.218.000.02
code18.000.343.000.00
reference9.000.111.000.01
Four sources and four numbers each. A reader can tell from this that the corpus is predominantly crawl, that the crawl was heavily deduplicated, and that nothing has been checked for overlap in the reference set, which is the one place to look first.

The figures that carry information are these. Composition, as the share of

tokens from each source, which is the single most predictive number about model

behaviour. Filter survival, the fraction of raw bytes that reached the final

corpus, which says how aggressive the pipeline was. Duplication removed, by

source, which says whether the deduplication actually ran. Benchmark overlap, by

benchmark, which decides whether the evaluations mean anything. Document length

distribution, not the mean, because the mean hides a tail of ten-word pages.

Language shares. Date range, because a corpus with no documents after a certain

year cannot know anything that happened later. And the labelled portion

separately: how many items, labelled by whom, under what instructions, at what

agreement.

FIG 2Where the tokens came from
The same corpus as a composition. This one chart predicts more about the finished model than the total does, and it is the first thing a reader should be shown.

The description is a document

Each of those numbers is nearly free while the pipeline is running, because the

stage that computes it is already touching every document. Each is expensive or

impossible afterwards. Nobody re-reads two trillion tokens to find out what the

length distribution was, so the honest answer six months later is that it is not

known.

So the description is assembled as the corpus is built: each stage writes its own

counts as it finishes, and the document is the collection of them rather than a

study commissioned at the end. It is worth fixing the set of questions in advance

so that stages know what to record. The useful set is: what is in it, where each

part came from, under what licence or permission, what was removed and by which

rule, what was labelled and by whom, what overlap with evaluations was measured,

what the intended use is, and what uses the authors consider unsupported.

The lesson stops here

3 more paragraphs to go

You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.

The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.

See the planThe contents

This is the reading half

Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.

The contents

The rest of this course

  1. 01What the Data Fixes, and What It Cannot
  2. 02Five Places, and What Each One Is Biased Towardsopening only
  3. 03Filters, In the Order That Costs Leastopening only
  4. 04Deduplication Is the Highest-Value Hour in the Projectopening only
  5. 05Setting the Proportions on Purposeopening only
  6. 06The Benchmark That Lies in the Flattering Directionopening only
  7. 07Writing a Task Two People Can Agree Onopening only
  8. 08Describing a Corpus in Numbers a Reader Can Checkyou are here

Read alongside