The library

How models work

Be able to explain how a model can hold far more parameters than it uses on any one input, how the choice of which parts to use is made and trained, what goes wrong with that choice, and when the arrangement is worth its costs

Many Specialists, One Model

A model does not have to run all of itself on every input. It can hold many parts and use a few, which breaks the link between how much it knows and what it costs to run.

8 lessons, written and corrected before you arrived. Reading them here needs no account. The first reads the whole way through; the others open and then stop, because a page nobody owns cannot tell who is reading it. Starting the course gives you your own copy, where every idea has problems standing under it and you can ask about any sentence.

Start reading

  1. 01Knowing More Without Doing MoreIn an ordinary model every parameter runs on every input, so what it knows and what it costs are the same number. Replacing one block with many and visiting a few separates them.
  2. 02One Small Matrix Decides Everythingopening onlyA single matrix scores every copy for every token, the highest few are visited, and their outputs are blended by those same scores. That last detail is the only reason the chooser learns anything.
  3. 03Not Chemistry, Not Law, Not Medicineopening onlyThe copies do specialise, but almost never by subject. They divide up surface properties of single tokens, and the arrangements that work best are the ones that make that division easier to express.
  4. 04The Favourite That Makes Itselfopening onlyA copy that gets slightly more traffic early gets better, which earns it more traffic, which makes it better still. Left alone the arrangement collapses onto a handful of copies and wastes the rest.
  5. 05Some Tokens Do Not Get Inopening onlyHardware wants fixed shapes, so each copy gets a fixed buffer. Tokens beyond it are silently skipped, and unused space is computed on anyway as though it were data.
  6. 06The Exponential in the Middle of Everythingopening onlySparse models diverge more readily than dense ones, and the reason sits in the chooser: an exponential of unbounded scores, computed in low precision, deciding things discretely.
  7. 07The Model You Must Hold and the Model You Runopening onlyServing a sparse model means holding every copy on hardware that will use two of them per token, moving tokens between machines twice per layer, and needing a large batch before any of it pays.
  8. 08The Comparison That Actually Means Somethingopening onlySparsity is a third dial alongside size and data, and whether turning it up helps depends on what you compare against, what resource is scarce, and how many requests arrive at once.