ContentsThe library

Judging a System, Not a Model

Four Reasons Your Number Disagrees With Your Users

Last timeCatching It Before a User Does

The offline score went up and the users got unhappier. That gap has four usual causes, all of them findable, and none of them a reason to stop measuring.

The four causes

Here is the situation this course exists for. The weekly paired comparison

favoured the new version. The nightly score went up two points. Support

tickets are up a third and the people who use the product most are saying it

got worse.

The two wrong responses are both common. One is to trust the number and

explain the complaints away as a vocal minority. The other is to abandon the

evaluation, which hands every future decision to whoever argues best.

The right response is to treat the gap as a finding. Something is true about

your evaluation that was not true yesterday, and finding out what is a day of

work.

FIG 1Four causes, four tests
per cent of caseshours to test for itfixable without new infr
the set no longer matche3821
the rubric wants somethi2441
the evaluation never saw2260
an aggregate hid a group1611
Indicative proportions from incident reviews where an offline score and user behaviour disagreed. Run the tests in order of cost: the breakdown by group takes an hour and explains a sixth of cases, so it goes first despite being the least common.

The set has drifted

The cheapest test after the breakdown, and the most frequent cause.

Take last week of traffic, classify it into the same groups your set uses, and

compare the proportions. Then look at what the set does not contain at all.

Both numbers tend to be surprising, because traffic moves for reasons that

have nothing to do with the product team: a new feature shipped, a partner

started sending requests, a user community found the thing and started using

it in a way nobody designed for.

FIG 2One diagnosis, in order
steptesthoursresultconclusionwhat happened
1break the score down by group1every group up or flatnot the aggregate hiding somethingAn hour, and it rules out the sharpest of the four causes. Always first.
2compare set composition to last week tra2follow-up questions are 22 per cent of tcandidate foundA category that barely exists in the set has become nearly a quarter of what users send. The evaluation is close to blind to it.
3mark fifty of those cases by hand4the new version is clearly worse on themconfirmedDirect confirmation on the suspected slice. Four hours and the question is settled.
4add a quota for them and re-run2the paired comparison now favours the olthe evaluation now sees itThe repair is to the set, not to the number. From here the original decision can be made again on evidence.
4 steps
Nine hours, in order of cost, each step either ruling out a cause or pointing at the next test. The last row is the part that matters: the outcome of a diagnosis is a better evaluation, not an explanation of why the old one was right.

Measuring something users do not experience

The third cause is the one that resists the hardest, because the evaluation

can be entirely correct and still miss the point.

Your evaluation reads an answer and judges its content. A user experiences an

answer arriving after some delay, formatted some way, on a screen of some size,

in the middle of a conversation, with an interface around it. Any of those can

get worse while content gets better, and your evaluation sees none of them.

The usual specific culprits: latency, when a change that improves answers also

doubles the time to first word. Formatting, when answers become long structured

lists that read well in a terminal and badly on a phone. Conversation, when

each answer is better alone and the system has stopped tracking what was said

two messages ago, which no single-turn evaluation can see. And refusal rate,

where a change that made the system more careful also made it decline things it

used to do.

The lesson stops here

3 more paragraphs to go

You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.

The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.

See the planThe contents

This is the reading half

Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.

The contents

The rest of this course

  1. 01The Boundary Decides What the Numbers Mean
  2. 02The Set Is a Sample, and You Chose the Samplingopening only
  3. 03If Two People Score It Differently, You Have No Measurementopening only
  4. 04A Marker That Never Gets Tired and Has Four Known Habitsopening only
  5. 05Three Points on Four Hundred Cases Is Nothingopening only
  6. 06Run Both on the Same Cases and Most of the Noise Disappearsopening only
  7. 07A Small Fast Set That Blocks, and Two Slower Ones That Do Notopening only
  8. 08Four Reasons Your Number Disagrees With Your Usersyou are here

Read alongside