Where Training Data Comes From
Filters, In the Order That Costs Least
Last timeWhere the Text Comes From
Most of a crawl is discarded, and the order the filters run in decides what the whole pipeline costs. The reasonable-looking filters are the dangerous ones.
Everything is a filter, and they are not the same price
Between a fetched page and a training token sit perhaps a dozen decisions to
discard something. They range from trivially cheap to genuinely expensive, and
the spread is larger than people expect.
Checking whether a document is long enough costs a comparison. Detecting its
language costs a small model over the first few hundred characters. Computing a
repetition statistic costs a pass over the text. Scoring it with a learned
quality model costs a neural network forward pass. Comparing it against every
other document for near-duplication costs a hashing scheme and a great deal of
memory.
| step | stage | documents in, millions | kept, per cent | cost per document, microseconds | what happened |
|---|---|---|---|---|---|
| 1 | 1 | 1000 | 62 | 2 | Length and character checks. Removes empty pages, navigation stubs and anything under a hundred words. |
| 2 | 2 | 620 | 71 | 40 | Language detection. Removes what is not the target language, and a fair amount of code and tables misread as prose. |
| 3 | 3 | 440 | 84 | 90 | Repetition statistics. Removes generated listings, keyword stuffing and pages that are one sentence repeated. |
| 4 | 4 | 370 | 73 | 120 | Blocklists and pattern rules. The cheapest stage to write and the one with the most collateral damage. |
| 5 | 5 | 270 | 66 | 4000 | Learned quality score. Expensive, so it runs last, on a quarter of what arrived. |
| 6 | 6 | 178 | 100 | 0 | What goes forward: 178 million of a billion, which is under a fifth. |
The ordering rule is simple: a filter should run before any filter more
expensive than itself, because every document it removes is one the expensive
stage never sees. Since the cheap filters are also the ones removing the most
obvious rubbish, the ordering that costs least is usually also the ordering that
works best.
Survival compounds
Each stage above sounds reasonable in isolation. Together they remove four
documents in five, and nobody decided that.
The lesson stops here
3 more paragraphs to go
You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.
The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.
See the planThe contentsThis is the reading half
Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.
The contents