Where Training Data Comes From
Setting the Proportions on Purpose
Last timeThe Same Thing Fifty Times
A mixture is a set of numbers that sum to one, so raising any source lowers every other. Here is how those numbers are chosen and what each choice costs.
There is no adding
A mixture is a list of sources with a share each, and the shares sum to one.
That sounds like a triviality and it is the single most commonly forgotten fact
in corpus work.
It means there is no operation called "adding more code". There is only
"raising code from 13 per cent to 20 per cent and reducing everything else by
eight per cent of itself". The second description is the true one, and it names
the cost. Somebody asking for more of their favourite source is, whether they
know it or not, asking for less of every other, and the conversation goes
better when the request is stated that way.
- the share of the mixture drawn from source i
- the weight given to that source, in tokens after any repetition
- the total over all sources, which is what makes the shares a proportion
| step | source | share before | share after | change, per cent of itself | what happened |
|---|---|---|---|---|---|
| 1 | 1 | 13 | 20 | 54 | Code, raised deliberately from 13 to 20 per cent. This is the change that was asked for. |
| 2 | 2 | 67 | 61.6 | -8 | Web, reduced by eight per cent of itself. Nobody decided this and it is the largest absolute change in the table. |
| 3 | 3 | 8 | 7.4 | -8 | Books, down by the same proportion. |
| 4 | 4 | 4 | 3.7 | -8 | Papers, down the same. At this size the loss may be enough to matter for a capability that was already marginal. |
Repeating a small source
Curated sources are small. A good reference corpus might be six billion tokens
against a mixture of 1.2 trillion, which is half a per cent if used once. To
give it six per cent you repeat it twelve times.
This is standard, deliberate and quite different from the accidental
duplication of the previous topic, because it is chosen, counted and written
down. But it is the same mechanism, and it carries the same consequence: the
repeat count is also the memorisation dial. Twelve passes over the same text is
exactly the regime in which a model begins to reproduce passages verbatim.
So the repeat count is a trade rather than a free choice. The usual resolutions
are to cap it somewhere in the low single digits for anything containing
personal or sensitive material, to accept a higher count for text that is
public, stable and that you would not mind being recited, and to prefer getting
more distinct text over repeating what you have whenever that is possible at
all.
Reweighting by a power
The third tool addresses a specific problem: when sources differ in size by a
factor of a thousand, sampling in proportion to size means the small ones never
appear, and sampling them equally means the large one is wasted.
The standard answer is to raise every share to a power below one and
renormalise.
The lesson stops here
3 more paragraphs to go
You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.
The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.
See the planThe contentsThis is the reading half
Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.
The contents