ContentsThe library

Judging a System, Not a Model

If Two People Score It Differently, You Have No Measurement

Last timeChoosing the Cases

Most evaluations fail here: the rule for what counts as a good answer was never written, so every score is partly a measurement of who happened to mark it.

The rule two people share

Here is a test that takes an afternoon and that very few teams have run. Take

thirty answers from your system. Give them to two colleagues separately, with

whatever instructions your evaluation currently uses, and ask each to score

them. Then compare.

The usual result is that they agree on about two thirds, which sounds

reasonable and is not. If both markers call three quarters of answers

acceptable, then two people answering at random in those proportions would

agree on 62 per cent of items without reading anything. Two thirds is barely

above guessing.

FIG 1The same answer, two markers, before and after a rubric
stepmarker Amarker Bagreedwhat they were givenwhat happened
1goodpoornoscore the answer for helpfulnessMarker A rewarded the thorough explanation. Marker B marked it down for not answering until the third paragraph. Both are defensible readings of the word helpful.
22 of 32 of 3yesthree checks: answers the question, statSame answer, same two people. The checklist removes the question of what helpful means and replaces it with three questions about the text.
2 steps
Nothing about the answer changed between the rows. What changed is that the second instruction can be applied without the marker supplying a definition of their own, which is the entire difference between a measurement and a collection of opinions.

What the test gives you is not a verdict on your colleagues. It is a

measurement of your instructions, and the instructions are the part of an

evaluation that almost never gets reviewed.

Checklists instead of adjectives

The repair is mechanical. Take the adjective and ask what observable fact about

the text would make you apply it.

Helpful becomes: does it answer the question that was asked, rather than a

nearby one. Does it state the limitation if there is one. Does it give the

reader something to do next. Three yes-or-no questions about the words on the

page, none of which requires a theory of helpfulness.

The lesson stops here

5 more paragraphs to go

You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.

The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.

See the planThe contents

This is the reading half

Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.

The contents

The rest of this course

  1. 01The Boundary Decides What the Numbers Mean
  2. 02The Set Is a Sample, and You Chose the Samplingopening only
  3. 03If Two People Score It Differently, You Have No Measurementyou are here
  4. 04A Marker That Never Gets Tired and Has Four Known Habitsopening only
  5. 05Three Points on Four Hundred Cases Is Nothingopening only
  6. 06Run Both on the Same Cases and Most of the Noise Disappearsopening only
  7. 07A Small Fast Set That Blocks, and Two Slower Ones That Do Notopening only
  8. 08Four Reasons Your Number Disagrees With Your Usersopening only

Read alongside