The library

Putting a model into service

Shrink a model deliberately rather than hopefully, well enough to predict the memory, the speed and the quality you will get from each technique before you run it, and to say which of them compose and which fight

Making a Model Smaller

A model that will not fit, or costs too much to run, can be made smaller in four quite different ways. This course derives what each one actually does to the numbers, what it costs in quality, and which combinations are safe.

8 lessons, written and corrected before you arrived. Reading them here needs no account. The first reads the whole way through; the others open and then stop, because a page nobody owns cannot tell who is reading it. Starting the course gives you your own copy, where every idea has problems standing under it and you can ask about any sentence.

Start reading

  1. 01Three Budgets, and Only One of Them Is the WeightsA served model occupies memory in three quite separate ways, and the famous techniques each attack only one of them. Here is the accounting, done in bytes.
  2. 02Round It, Store the Integer, Multiply It Backopening onlyStoring a weight in four bits is two multiplications and a rounding. Doing the arithmetic on paper is what makes the failure case obvious before it reaches your model.
  3. 03Most of the Model Does Not Care, and You Have to Find the Part That Doesopening onlyRounding damage is not spread evenly. A few per cent of a model carries most of the sensitivity, and keeping that part precise costs almost nothing.
  4. 04Round the Finished Model, or Train One That Expects Itopening onlyRounding a trained model takes an hour and a few hundred sample inputs. Training one that knows it will be rounded costs a run and buys about two bits.
  5. 05A Zero Still Occupies Its Place in the Rowopening onlyRemoving scattered weights and removing whole rows sound like the same technique. Only one of them makes a model smaller or faster on ordinary hardware.
  6. 06The Wrong Answers Are Where the Teaching Isopening onlyA small model trained on a large one's full output distribution learns more than the same model trained on the original labels. Here is the objective and why it works.
  7. 07The Average Moved Half a Point and the Model Cannot Add Any Moreopening onlyCompression damage is concentrated, not spread. An average score is the one measurement guaranteed to miss it, and here is what to measure instead.
  8. 08The Savings Multiply and So Does the Damageopening onlyCompression techniques combine, but not additively: the memory savings multiply and the quality costs more than add. Order decides how much more, and whether you can tell what did it.