Where Training Data Comes From
Deduplication Is the Highest-Value Hour in the Project
Last timeThrowing Most of It Away
Duplicate text wastes compute, changes what the model memorises, and does not show up in the loss. Finding near-duplicates at scale is a solved problem worth knowing.
Two separate harms
Web corpora contain a great deal of repeated text. The same news article is
syndicated to forty sites. The same licence text appears in a hundred thousand
code repositories. The same forum thread is mirrored, scraped and republished.
The same product description appears on every reseller.
This causes two different problems, and they are worth separating because they
have different consequences.
The first is waste. A passage appearing fifty times is seen fifty times during
one pass over the corpus. The model spends fifty steps' worth of compute on it
and learns what one would have taught. Measured against the corpus size, a
third of the compute of a typical unfiltered run buys nothing.
The second is memorisation. The probability that a model can reproduce a passage
word for word, given a short prompt from it, rises sharply with the number of
times that passage appeared. This is how private information, licence keys and
personal details leak out of models, and the mechanism is repetition rather than
anything exotic.
| step | copies kept | steps spent on it | distinct text learned | verbatim recall | what happened |
|---|---|---|---|---|---|
| 1 | 41 | 41 | 1 | 1 | As crawled. Forty-one steps of compute, one article's worth of information, and the model can recite it. |
| 2 | 8 | 8 | 1 | 1 | Partial deduplication. Still reciting it, because the threshold is lower than people assume. |
| 3 | 1 | 1 | 1 | 0 | One copy. The same information, a fortieth of the compute, and no verbatim recall. |
Neither harm appears in the loss, and the reason is worth stating. The held-out
set is drawn from the same corpus, so it contains the same duplicates, and a
model that has memorised them scores well on them. A duplicated corpus produces
a flattering loss, which is the worst possible combination: a cost you are
paying and a measurement that rewards you for it.
The duplicates are not exact
Exact duplicate detection is trivial and finds almost nothing, because the
copies are never quite the same.
The syndicated article has a different site name in the header, a different
date format, a different cookie notice and a different set of related-article
links. The mirrored forum thread has a different advertisement block. The
republished documentation has a different version number. Every one of these
differences is outside the text anybody cares about, and every one of them
defeats an exact comparison.
So the real question is similarity, and the standard measure compares the sets
of short word sequences each document contains.
The lesson stops here
4 more paragraphs to go
You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.
The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.
See the planThe contentsThis is the reading half
Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.
The contents