ContentsThe library

Where Training Data Comes From

Writing a Task Two People Can Agree On

Last timeWhen the Test Is in the Training

Human-labelled data is small, expensive and decisive. Almost all of its quality comes from the instructions, and agreement between labellers is how you find out.

The instructions are the work

Most of a corpus is found. A small part of it has to be written by people, and

that small part decides a disproportionate amount: it is where the model learns

what a good answer looks like, what to refuse, and what the task even is.

When a labelled set comes back inconsistent, the instinctive diagnosis is that

the labellers were careless. That is occasionally true and usually wrong. The

far more common cause is that the task was underspecified, and two conscientious

people reading the same instruction reached different conclusions because the

instruction did not cover the case in front of them.

A labelling task is, formally, a function from a piece of text to a label. The

instructions are the only place that function is written down. An instruction

that says "mark the response as helpful or unhelpful" has not defined a

function; it has named one and left it undefined on almost every input. What

about a response that is helpful but wrong? Helpful but too long? Correct but

answering a question that was not asked? Each of those is a case the labellers

will meet in the first hour.

FIG 1Three rounds of the same task
steproundlabelled itemsraw agreement, per centchance-corrected agreementwhat happened
11200710.31One paragraph of instructions. Corrected agreement of 0.31 is barely better than guessing, and the raw 71 per cent hides that.
22200840.63After reading the disagreements and writing rules for the six cases they revealed. The instructions are now four pages.
33200910.8After one more pass. 0.80 is the usual threshold for a task considered well specified.
3 steps
Six hundred items of labelling, almost all of which was spent learning what the task was. This is the normal shape of a labelling project and it should be budgeted for.

The practical procedure is a loop. Write the instructions. Have two people label

a small batch independently. Read every disagreement. For each one, decide what

the right answer is and write the rule that produces it. Repeat until agreement

stops improving. Only then label at volume.

Agreement, corrected for chance

Raw agreement is a misleading number whenever one label is more common than the

others, which is almost always.

Suppose 85 per cent of responses are acceptable and 15 per cent are not. Two

labellers who mark everything acceptable without reading agree 85 per cent of

the time. The number looks excellent and no labelling has occurred. The fix is

to subtract the agreement chance alone would produce.

The lesson stops here

4 more paragraphs to go

You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.

The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.

See the planThe contents

This is the reading half

Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.

The contents

The rest of this course

  1. 01What the Data Fixes, and What It Cannot
  2. 02Five Places, and What Each One Is Biased Towardsopening only
  3. 03Filters, In the Order That Costs Leastopening only
  4. 04Deduplication Is the Highest-Value Hour in the Projectopening only
  5. 05Setting the Proportions on Purposeopening only
  6. 06The Benchmark That Lies in the Flattering Directionopening only
  7. 07Writing a Task Two People Can Agree Onyou are here
  8. 08Describing a Corpus in Numbers a Reader Can Checkopening only

Read alongside