The library

How models work

Follow one model from a pile of text to a set of weights that work, naming every decision taken on the way and what each one costs, well enough to read somebody else's training report and know which choice explains the number you are looking at

Training a Model From Scratch

Every other course here takes a trained model as given. This one does the training: the loop, where the data comes from, where the weights start, how fast to move, and how to tell a run that is working from one that is quietly dead.

8 lessons, written and corrected before you arrived. Reading them here needs no account. The first reads the whole way through; the others open and then stop, because a page nobody owns cannot tell who is reading it. Starting the course gives you your own copy, where every idea has problems standing under it and you can ask about any sentence.

Start reading

  1. 01Four Lines, and Genuinely All of ItTraining is one short loop repeated: predict, score, measure the blame, move the weights. Everything hard about training lives outside those four lines.
  2. 02The Four Hops Between Storage and the Stepopening onlyA batch travels from a store, through host memory, through tokenisation, onto the device. Each hop can be the one the whole run is waiting on.
  3. 03The Starting Scale, and Why It Decides Everythingopening onlyWeights start as random numbers, and the one thing that matters about them is their scale: get it wrong and a deep network cannot be trained at all.
  4. 04One Number, and Why It Cannot Be One Numberopening onlyThe learning rate changes the outcome of a run more than any architectural choice, and the right value at step one is the wrong value at step fifty thousand.
  5. 05What a Bigger Batch Actually Buysopening onlyBatch size sets how noisy each gradient is, and the noise falls only with the square root, which is why doubling a batch buys far less than it costs.
  6. 06The Four Numbers Worth Watchingopening onlyA run reports hundreds of numbers and four of them matter: training loss, held-out loss, gradient size and throughput. Each fails in its own way.
  7. 07Three Ways a Run Breaks, and What Each One Looks Likeopening onlyA loss spike, a divergence and a dead plateau look alike for about ten steps and need completely different responses. Here is how to tell them apart.
  8. 08Nobody Trains to Convergenceopening onlyRuns end because the budget ran out, not because the model converged. Deciding what to keep, and on what evidence, is a separate skill from training.