The library

The systems around the model

Spread one training run over many devices, well enough to choose the split from the shape of the model and the speed of the link, to predict the efficiency you will get, and to find where a slow run is actually waiting

Training Across Many Machines

One machine stops being enough long before a model is interesting. This course covers the ways a run can be split across devices, what each one sends over the wire, and why the fast link between chips decides the whole design.

8 lessons, written and corrected before you arrived. Reading them here needs no account. The first reads the whole way through; the others open and then stop, because a page nobody owns cannot tell who is reading it. Starting the course gives you your own copy, where every idea has problems standing under it and you can ask about any sentence.

Start reading

  1. 01The Weights Are the Smallest Thing You Have to StoreTraining needs several copies of every parameter, not one. Counting them is how you find out whether a run fits on one device, and which budget runs out first.
  2. 02Every Machine Does the Whole Thing, and Then They Agreeopening onlyThe simplest split gives every device the whole model and a different slice of the batch. What they have to exchange each step is fixed, and that fact decides the efficiency.
  3. 03Nobody Is in Charge and That Is Why It Scalesopening onlySumming a value across every device looks like it should cost more as devices are added. Arranged as a ring it does not, and the derivation is the most useful thing in this course.
  4. 04Half a Layer Each, and a Phone Call in the Middle of Every Oneopening onlySplitting a single layer across devices means exchanging numbers partway through the forward pass. That traffic scales with tokens, not parameters, which puts it on the fastest wire you own.
  5. 05Four Machines, and Three of Them Are Waitingopening onlyLay the layers out in stages and each device waits for the one before it. The idle fraction follows from two numbers, and splitting the batch into pieces is what shrinks it.
  6. 06Sixteen Copies of the Same Numbers, in Sixteen Placesopening onlyThe data split has every device storing an identical copy of everything. Stop doing that, hand each device a slice, and fetch the rest just before it is needed.
  7. 07You Do Not Choose One, You Multiply Fouropening onlyThe four ways of dividing training are not alternatives. They compose, and each one is assigned to the link whose speed its traffic can live with.
  8. 08Three Hundred Devices Waiting on Oneopening onlyPoor scaling has three usual causes and they leave different marks on a timeline. Read the marks and you know which one it is without changing anything.