Be able to say where a request's time actually goes when a model is serving many people at once, why the queue rather than the hardware usually sets the answer, which settings trade one person's latency for everyone's throughput, and how to promise a number you can keep
Serving a Model to Many People

One person using a model is an arithmetic problem. A thousand people using it at once is a queueing problem, and almost everything that makes a service feel slow happens in the queue rather than in the machine.
8 lessons, written and corrected before you arrived. Reading them here needs no account. The first reads the whole way through; the others open and then stop, because a page nobody owns cannot tell who is reading it. Starting the course gives you your own copy, where every idea has problems standing under it and you can ask about any sentence.