ContentsThe library

Judging a System, Not a Model

A Small Fast Set That Blocks, and Two Slower Ones That Do Not

Last timeComparing Two Versions

The evaluation that catches regressions is not the thorough one. It is the two-minute one that runs on every change and refuses to let three specific things through.

The set that runs on everything

Everything so far has been about measuring well. This lesson is about

measuring fast, which is a different goal and is often better served by a worse

evaluation.

A thorough run takes an hour and costs real money, so it happens nightly at

best. That means a regression introduced at ten in the morning reaches users

and stays with them until the following day. The fix is not to make the

thorough run faster. It is to put something small in front of it.

FIG 1Three tiers, by speed
The split is by speed rather than by importance. The most important measurement is the weekly one, and it is also the one furthest from the keyboard, because a measurement that takes an hour cannot be in the path of every change without stopping the team.

The fast set is forty to eighty cases, all with closed answers that can be

checked automatically, chosen on one principle: each one broke at some point.

A regression suite grows out of incidents, not out of coverage planning. When

something reaches users and is fixed, the case that revealed it joins the set,

with a note saying what it is guarding.

FIG 2The three tiers
casesminutes per runruns per weekblocks a merge
fast set60.02.0300.01.0
nightly full set600.055.01.00.0
weekly paired comparison400.0180.00.20.0
One row blocks. The fast set runs three hundred times a week and has to stay fast enough that nobody minds, which is the constraint that sets its size. The other two are allowed to be slow because nothing waits on them.

What it should block on

The temptation is to put the score in the gate: block if the score drops more

than two points. This is wrong for a reason the fifth lesson established.

On sixty cases the error on a score is over six points, so a two-point drop is

noise, and a gate wired to noise fails about as often as it succeeds.

Block on absolutes instead.

A case that passed on the previous commit and fails now. That is a specific,

reproducible, nameable failure, and it is the single most valuable condition in

the gate. Any crash or timeout. An empty or truncated answer. A safety check

failing, which is absolute by nature. A structural violation, such as output

that is supposed to be a particular shape and is not.

Every one of those is a yes or no about a specific case, which means the

failure message can say which case and what happened, and the person who

caused it can fix it in minutes rather than opening an investigation.

The lesson stops here

3 more paragraphs to go

You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.

The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.

See the planThe contents

This is the reading half

Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.

The contents

The rest of this course

  1. 01The Boundary Decides What the Numbers Mean
  2. 02The Set Is a Sample, and You Chose the Samplingopening only
  3. 03If Two People Score It Differently, You Have No Measurementopening only
  4. 04A Marker That Never Gets Tired and Has Four Known Habitsopening only
  5. 05Three Points on Four Hundred Cases Is Nothingopening only
  6. 06Run Both on the Same Cases and Most of the Noise Disappearsopening only
  7. 07A Small Fast Set That Blocks, and Two Slower Ones That Do Notyou are here
  8. 08Four Reasons Your Number Disagrees With Your Usersopening only

Read alongside