The library

Putting a model into service

Build an evaluation for a product that has a model inside it, well enough to tell a real improvement from noise, to catch a regression before a user does, and to say honestly what your numbers do not cover

Judging a System, Not a Model

Benchmarks score models. Nobody ships a model, they ship a system with retrieval, prompts, tools and fallbacks in it. This course builds the evaluation for that: the set, the judge, the error bars, and what to do when it disagrees with your users.

8 lessons, written and corrected before you arrived. Reading them here needs no account. The first reads the whole way through; the others open and then stop, because a page nobody owns cannot tell who is reading it. Starting the course gives you your own copy, where every idea has problems standing under it and you can ask about any sentence.

Start reading

  1. 01The Boundary Decides What the Numbers MeanAn evaluation measures whatever sits inside the line you drew, and most teams never draw it. Here is where the line goes and what has to be pinned once it is drawn.
  2. 02The Set Is a Sample, and You Chose the Samplingopening onlyAn evaluation set is a claim about which inputs matter. Build it from your own traffic, state the sampling behind each group, and keep part of it unseen.
  3. 03If Two People Score It Differently, You Have No Measurementopening onlyMost evaluations fail here: the rule for what counts as a good answer was never written, so every score is partly a measurement of who happened to mark it.
  4. 04A Marker That Never Gets Tired and Has Four Known Habitsopening onlyA model can apply your rubric at a thousandth of the cost of a person. First measure how often it agrees with people, then handle the four biases it brings to the job.
  5. 05Three Points on Four Hundred Cases Is Nothingopening onlyA score without an error bar cannot answer the only question anybody asks of it. Here is the arithmetic, and the smallest difference your set can honestly detect.
  6. 06Run Both on the Same Cases and Most of the Noise Disappearsopening onlyThe single technique that buys the most in evaluation: compare per case rather than comparing two averages, and the hard cases stop counting against you.
  7. 07A Small Fast Set That Blocks, and Two Slower Ones That Do Notopening onlyThe evaluation that catches regressions is not the thorough one. It is the two-minute one that runs on every change and refuses to let three specific things through.
  8. 08Four Reasons Your Number Disagrees With Your Usersopening onlyThe offline score went up and the users got unhappier. That gap has four usual causes, all of them findable, and none of them a reason to stop measuring.