ContentsThe library

Where Training Data Comes From

Deduplication Is the Highest-Value Hour in the Project

Last timeThrowing Most of It Away

Duplicate text wastes compute, changes what the model memorises, and does not show up in the loss. Finding near-duplicates at scale is a solved problem worth knowing.

Two separate harms

Web corpora contain a great deal of repeated text. The same news article is

syndicated to forty sites. The same licence text appears in a hundred thousand

code repositories. The same forum thread is mirrored, scraped and republished.

The same product description appears on every reseller.

This causes two different problems, and they are worth separating because they

have different consequences.

The first is waste. A passage appearing fifty times is seen fifty times during

one pass over the corpus. The model spends fifty steps' worth of compute on it

and learns what one would have taught. Measured against the corpus size, a

third of the compute of a typical unfiltered run buys nothing.

The second is memorisation. The probability that a model can reproduce a passage

word for word, given a short prompt from it, rises sharply with the number of

times that passage appeared. This is how private information, licence keys and

personal details leak out of models, and the mechanism is repetition rather than

anything exotic.

FIG 1One article, forty-one copies
stepcopies keptsteps spent on itdistinct text learnedverbatim recallwhat happened
1414111As crawled. Forty-one steps of compute, one article's worth of information, and the model can recite it.
28811Partial deduplication. Still reciting it, because the threshold is lower than people assume.
31110One copy. The same information, a fortieth of the compute, and no verbatim recall.
3 steps
The third column never changes, which is the whole argument. Nothing is lost by removing the copies, and two separate things are gained.
FIG 2Where the duplication comes from
None of this is unusual or avoidable material. It is the ordinary structure of the web, which is why every crawl has it and why the fix has to be mechanical.

Neither harm appears in the loss, and the reason is worth stating. The held-out

set is drawn from the same corpus, so it contains the same duplicates, and a

model that has memorised them scores well on them. A duplicated corpus produces

a flattering loss, which is the worst possible combination: a cost you are

paying and a measurement that rewards you for it.

The duplicates are not exact

Exact duplicate detection is trivial and finds almost nothing, because the

copies are never quite the same.

The syndicated article has a different site name in the header, a different

date format, a different cookie notice and a different set of related-article

links. The mirrored forum thread has a different advertisement block. The

republished documentation has a different version number. Every one of these

differences is outside the text anybody cares about, and every one of them

defeats an exact comparison.

So the real question is similarity, and the standard measure compares the sets

of short word sequences each document contains.

The lesson stops here

4 more paragraphs to go

You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.

The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.

See the planThe contents

This is the reading half

Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.

The contents

The rest of this course

  1. 01What the Data Fixes, and What It Cannot
  2. 02Five Places, and What Each One Is Biased Towardsopening only
  3. 03Filters, In the Order That Costs Leastopening only
  4. 04Deduplication Is the Highest-Value Hour in the Projectyou are here
  5. 05Setting the Proportions on Purposeopening only
  6. 06The Benchmark That Lies in the Flattering Directionopening only
  7. 07Writing a Task Two People Can Agree Onopening only
  8. 08Describing a Corpus in Numbers a Reader Can Checkopening only

Read alongside