Four Reasons Your Number Disagrees With Your Users
Last timeCatching It Before a User Does
The offline score went up and the users got unhappier. That gap has four usual causes, all of them findable, and none of them a reason to stop measuring.
The four causes
Here is the situation this course exists for. The weekly paired comparison
favoured the new version. The nightly score went up two points. Support
tickets are up a third and the people who use the product most are saying it
got worse.
The two wrong responses are both common. One is to trust the number and
explain the complaints away as a vocal minority. The other is to abandon the
evaluation, which hands every future decision to whoever argues best.
The right response is to treat the gap as a finding. Something is true about
your evaluation that was not true yesterday, and finding out what is a day of
work.
| per cent of cases | hours to test for it | fixable without new infr | |
|---|---|---|---|
| the set no longer matche | 38 | 2 | 1 |
| the rubric wants somethi | 24 | 4 | 1 |
| the evaluation never saw | 22 | 6 | 0 |
| an aggregate hid a group | 16 | 1 | 1 |
The set has drifted
The cheapest test after the breakdown, and the most frequent cause.
Take last week of traffic, classify it into the same groups your set uses, and
compare the proportions. Then look at what the set does not contain at all.
Both numbers tend to be surprising, because traffic moves for reasons that
have nothing to do with the product team: a new feature shipped, a partner
started sending requests, a user community found the thing and started using
it in a way nobody designed for.
| step | test | hours | result | conclusion | what happened |
|---|---|---|---|---|---|
| 1 | break the score down by group | 1 | every group up or flat | not the aggregate hiding something | An hour, and it rules out the sharpest of the four causes. Always first. |
| 2 | compare set composition to last week tra | 2 | follow-up questions are 22 per cent of t | candidate found | A category that barely exists in the set has become nearly a quarter of what users send. The evaluation is close to blind to it. |
| 3 | mark fifty of those cases by hand | 4 | the new version is clearly worse on them | confirmed | Direct confirmation on the suspected slice. Four hours and the question is settled. |
| 4 | add a quota for them and re-run | 2 | the paired comparison now favours the ol | the evaluation now sees it | The repair is to the set, not to the number. From here the original decision can be made again on evidence. |
Measuring something users do not experience
The third cause is the one that resists the hardest, because the evaluation
can be entirely correct and still miss the point.
Your evaluation reads an answer and judges its content. A user experiences an
answer arriving after some delay, formatted some way, on a screen of some size,
in the middle of a conversation, with an interface around it. Any of those can
get worse while content gets better, and your evaluation sees none of them.
The usual specific culprits: latency, when a change that improves answers also
doubles the time to first word. Formatting, when answers become long structured
lists that read well in a terminal and badly on a phone. Conversation, when
each answer is better alone and the system has stopped tracking what was said
two messages ago, which no single-turn evaluation can see. And refusal rate,
where a change that made the system more careful also made it decline things it
used to do.
The lesson stops here
3 more paragraphs to go
You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.
The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.
See the planThe contentsThis is the reading half
Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.
The contents