The library

How models work

Build and judge a training set rather than accepting one, well enough to say what a model will be good at before it is trained, and to spot the three data faults that explain most disappointing results

Where Training Data Comes From

A model is mostly its data, and almost nobody looks at the data. This course covers collecting it, cleaning it, removing the duplicates that quietly ruin it, and the contamination that makes a benchmark lie to you.

8 lessons, written and corrected before you arrived. Reading them here needs no account. The first reads the whole way through; the others open and then stop, because a page nobody owns cannot tell who is reading it. Starting the course gives you your own copy, where every idea has problems standing under it and you can ask about any sentence.

Start reading

  1. 01What the Data Fixes, and What It CannotA training set sets a ceiling no later choice can lift. It also leaves several things entirely open, and confusing the two is how teams spend a year on the wrong problem.
  2. 02Five Places, and What Each One Is Biased Towardsopening onlyAlmost every large corpus is built from the same handful of sources. Each arrives with a systematic lean that survives filtering, so knowing them is knowing the model.
  3. 03Filters, In the Order That Costs Leastopening onlyMost of a crawl is discarded, and the order the filters run in decides what the whole pipeline costs. The reasonable-looking filters are the dangerous ones.
  4. 04Deduplication Is the Highest-Value Hour in the Projectopening onlyDuplicate text wastes compute, changes what the model memorises, and does not show up in the loss. Finding near-duplicates at scale is a solved problem worth knowing.
  5. 05Setting the Proportions on Purposeopening onlyA mixture is a set of numbers that sum to one, so raising any source lowers every other. Here is how those numbers are chosen and what each choice costs.
  6. 06The Benchmark That Lies in the Flattering Directionopening onlyIf the evaluation set leaked into training, every number you report is wrong upwards, which is the direction nobody questions. Measuring the overlap is the only defence.
  7. 07Writing a Task Two People Can Agree Onopening onlyHuman-labelled data is small, expensive and decisive. Almost all of its quality comes from the instructions, and agreement between labellers is how you find out.
  8. 08Describing a Corpus in Numbers a Reader Can Checkopening onlyA finished corpus needs a short, honest description: its size, its composition, how it was filtered, who labelled what, and which claims about the model those numbers do not support.