The library

Putting a model into service

Be able to say where a request's time actually goes when a model is serving many people at once, why the queue rather than the hardware usually sets the answer, which settings trade one person's latency for everyone's throughput, and how to promise a number you can keep

Serving a Model to Many People

One person using a model is an arithmetic problem. A thousand people using it at once is a queueing problem, and almost everything that makes a service feel slow happens in the queue rather than in the machine.

8 lessons, written and corrected before you arrived. Reading them here needs no account. The first reads the whole way through; the others open and then stop, because a page nobody owns cannot tell who is reading it. Starting the course gives you your own copy, where every idea has problems standing under it and you can ask about any sentence.

Start reading

  1. 01Mostly WaitingA request's total time is what it spent waiting plus what it spent being served, and under load the first term is several times the second while the hardware looks fine.
  2. 02The Average Minute Does Not Existopening onlyRequests arrive in clumps, and the clumping is what builds queues. A service sized for its average traffic will be overwhelmed by traffic whose average it was sized for.
  3. 03Reading the Weights Once for Everybodyopening onlyThe expensive part of producing a word is reading the model, and the reading serves everyone in the group at once. Group size is the single largest lever in serving, and it is a trade.
  4. 04The Two Jobs That Get in Each Other's Wayopening onlyReading a prompt and writing an answer are different kinds of work on the same hardware. Mixed together they interfere, and the interference is felt by whoever was already mid-sentence.
  5. 05Where the Complaints Actually Liveopening onlyNobody experiences an average. They experience their own request, and if their page needs twenty of them, they experience the slowest of twenty rather than the typical one.
  6. 06Saying No While You Still Canopening onlyA service that accepts everything offered to it serves nobody once the offer exceeds its capacity. Refusing work early is the only way to keep the accepted work fast.
  7. 07Help That Arrives Nine Minutes Lateopening onlyAdding a machine is a decision whose effect appears minutes later. Everything about automatic scaling follows from that delay, including why spare capacity has to be paid for.
  8. 08A Number You Can Be Held Toopening onlyA latency promise is only meaningful if it names a percentile, a window, a place of measurement and what counts as a failure. Then it becomes a design constraint rather than an aspiration.