ContentsThe library

Where Training Data Comes From

Five Places, and What Each One Is Biased Towards

Last timeThe Data Decides

Almost every large corpus is built from the same handful of sources. Each arrives with a systematic lean that survives filtering, so knowing them is knowing the model.

The crawl is not the web

Most of a large corpus by volume is a web crawl, and almost everybody thinks of

a crawl as a sample of the web. It is not.

A crawler starts from a list of addresses and follows links. It gets what is

linked to, reachable without signing in, served quickly enough to be worth

waiting for, and permitted by the site. Everything else is invisible to it: the

inside of every forum that requires an account, every document behind a

paywall, everything in a database that is only rendered on request, everything

published somewhere nobody links to, and an enormous amount of material that is

simply served too slowly.

FIG 1From a page to a training shard
Each arrow removes material, and the removals are not random. The shape of what survives is set here, long before any quality filter runs.

Text extraction deserves more respect than it gets. A fetched page is mostly not

the thing you want: menus, footers, cookie banners, related-article lists and

advertising copy, with the article threaded through it. Getting the article out

cleanly is a hard, boring engineering problem, and the difference between a good

extractor and a mediocre one shows up in the final model more clearly than most

architectural choices do.

The curated few

Alongside the crawl sit a handful of sources that are small, countable and

deliberately included.

The lesson stops here

5 more paragraphs to go

You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.

The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.

See the planThe contents

This is the reading half

Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.

The contents

The rest of this course

  1. 01What the Data Fixes, and What It Cannot
  2. 02Five Places, and What Each One Is Biased Towardsyou are here
  3. 03Filters, In the Order That Costs Leastopening only
  4. 04Deduplication Is the Highest-Value Hour in the Projectopening only
  5. 05Setting the Proportions on Purposeopening only
  6. 06The Benchmark That Lies in the Flattering Directionopening only
  7. 07Writing a Task Two People Can Agree Onopening only
  8. 08Describing a Corpus in Numbers a Reader Can Checkopening only

Read alongside