ContentsThe library

Where Training Data Comes From

Filters, In the Order That Costs Least

Last timeWhere the Text Comes From

Most of a crawl is discarded, and the order the filters run in decides what the whole pipeline costs. The reasonable-looking filters are the dangerous ones.

Everything is a filter, and they are not the same price

Between a fetched page and a training token sit perhaps a dozen decisions to

discard something. They range from trivially cheap to genuinely expensive, and

the spread is larger than people expect.

Checking whether a document is long enough costs a comparison. Detecting its

language costs a small model over the first few hundred characters. Computing a

repetition statistic costs a pass over the text. Scoring it with a learned

quality model costs a neural network forward pass. Comparing it against every

other document for near-duplication costs a hashing scheme and a great deal of

memory.

FIG 1One pipeline, in order
stepstagedocuments in, millionskept, per centcost per document, microsecondswhat happened
111000622Length and character checks. Removes empty pages, navigation stubs and anything under a hundred words.
226207140Language detection. Removes what is not the target language, and a fair amount of code and tables misread as prose.
334408490Repetition statistics. Removes generated listings, keyword stuffing and pages that are one sentence repeated.
4437073120Blocklists and pattern rules. The cheapest stage to write and the one with the most collateral damage.
55270664000Learned quality score. Expensive, so it runs last, on a quarter of what arrived.
661781000What goes forward: 178 million of a billion, which is under a fifth.
6 steps
The order is the point. Had the quality model run first it would have scored a billion documents instead of 270 million, at four milliseconds each, which is nearly four times the compute of everything else combined.

The ordering rule is simple: a filter should run before any filter more

expensive than itself, because every document it removes is one the expensive

stage never sees. Since the cheap filters are also the ones removing the most

obvious rubbish, the ordering that costs least is usually also the ordering that

works best.

Survival compounds

Each stage above sounds reasonable in isolation. Together they remove four

documents in five, and nobody decided that.

The lesson stops here

3 more paragraphs to go

You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.

The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.

See the planThe contents

This is the reading half

Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.

The contents

The rest of this course

  1. 01What the Data Fixes, and What It Cannot
  2. 02Five Places, and What Each One Is Biased Towardsopening only
  3. 03Filters, In the Order That Costs Leastyou are here
  4. 04Deduplication Is the Highest-Value Hour in the Projectopening only
  5. 05Setting the Proportions on Purposeopening only
  6. 06The Benchmark That Lies in the Flattering Directionopening only
  7. 07Writing a Task Two People Can Agree Onopening only
  8. 08Describing a Corpus in Numbers a Reader Can Checkopening only

Read alongside