ContentsThe library

Judging a System, Not a Model

A Marker That Never Gets Tired and Has Four Known Habits

Last timeDeciding What Counts as Right

A model can apply your rubric at a thousandth of the cost of a person. First measure how often it agrees with people, then handle the four biases it brings to the job.

Why hand it to a model

The previous lesson ended with an uncomfortable number: six hundred open cases

at forty seconds each is most of a working day for one run. That cost decides

the cadence of your evaluation, and a monthly cadence means a regression lives

for up to a month.

A model applying the same checklist does it in minutes for a few pounds. The

argument for a model marker is not that it is as good as a person. It is that

it is good enough and frequent enough, and that a slightly noisier number

available on every change beats a better number available after the damage.

FIG 1What the two marking methods cost over a year of runs
0.0075.00150.00225.00300.000.015.030.045.060.0runs over the year
a person marks every runa model marks, built once and checked quarterly
The flat part of the second line is the work of writing the judge rule and measuring its agreement, which is real and happens once. After about three runs the model marker is cheaper, and by thirty it is cheaper by a factor of ten. That factor is what buys a daily cadence.

Measuring the marker

Before any of that, the judge is a system and has to be tested like one.

Take a sample, two hundred cases is usually enough, and mark it both ways: by

the model and by a person using the same checklist. Then compute how often they

agree, corrected for chance in the way the previous lesson described. Report

that figure next to every score the judge produces.

FIG 2Where the judge sits
The lower path is the part teams skip. A judge installed without it produces numbers that look identical to measured ones and carry no information about their own reliability.

Two cases need distinguishing. If agreement is high, the judge can mark

everything and people spot-check. If it is high on some checks and low on

others, which is the common result, let the judge mark the checks it is good at

and route the rest to a person. Factual checks tend to agree well. Checks

involving tone or appropriateness tend not to.

Re-measure when the judge model version changes. A judge upgraded quietly is a

measuring instrument recalibrated without telling anybody, and every comparison

across that boundary is invalid.

The lesson stops here

4 more paragraphs to go

You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.

The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.

See the planThe contents

This is the reading half

Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.

The contents

The rest of this course

  1. 01The Boundary Decides What the Numbers Mean
  2. 02The Set Is a Sample, and You Chose the Samplingopening only
  3. 03If Two People Score It Differently, You Have No Measurementopening only
  4. 04A Marker That Never Gets Tired and Has Four Known Habitsyou are here
  5. 05Three Points on Four Hundred Cases Is Nothingopening only
  6. 06Run Both on the Same Cases and Most of the Noise Disappearsopening only
  7. 07A Small Fast Set That Blocks, and Two Slower Ones That Do Notopening only
  8. 08Four Reasons Your Number Disagrees With Your Usersopening only

Read alongside