Where Training Data Comes From
Five Places, and What Each One Is Biased Towards
Last timeThe Data Decides
Almost every large corpus is built from the same handful of sources. Each arrives with a systematic lean that survives filtering, so knowing them is knowing the model.
The crawl is not the web
Most of a large corpus by volume is a web crawl, and almost everybody thinks of
a crawl as a sample of the web. It is not.
A crawler starts from a list of addresses and follows links. It gets what is
linked to, reachable without signing in, served quickly enough to be worth
waiting for, and permitted by the site. Everything else is invisible to it: the
inside of every forum that requires an account, every document behind a
paywall, everything in a database that is only rendered on request, everything
published somewhere nobody links to, and an enormous amount of material that is
simply served too slowly.
Text extraction deserves more respect than it gets. A fetched page is mostly not
the thing you want: menus, footers, cookie banners, related-article lists and
advertising copy, with the article threaded through it. Getting the article out
cleanly is a hard, boring engineering problem, and the difference between a good
extractor and a mediocre one shows up in the final model more clearly than most
architectural choices do.
The curated few
Alongside the crawl sit a handful of sources that are small, countable and
deliberately included.
The lesson stops here
5 more paragraphs to go
You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.
The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.
See the planThe contentsThis is the reading half
Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.
The contents