Build an evaluation for a product that has a model inside it, well enough to tell a real improvement from noise, to catch a regression before a user does, and to say honestly what your numbers do not cover
Judging a System, Not a Model

Benchmarks score models. Nobody ships a model, they ship a system with retrieval, prompts, tools and fallbacks in it. This course builds the evaluation for that: the set, the judge, the error bars, and what to do when it disagrees with your users.
8 lessons, written and corrected before you arrived. Reading them here needs no account. The first reads the whole way through; the others open and then stop, because a page nobody owns cannot tell who is reading it. Starting the course gives you your own copy, where every idea has problems standing under it and you can ask about any sentence.