ContentsThe library

Learning From Reward

A Reward Nobody Wrote, and the Tether That Keeps It Honest

Last timeKeeping the Step Small Where It Matters

A language model has no score to collect, so the reward is fitted from human comparisons, and because a fitted reward can be gamed the objective carries a tether to the model it started from.

Every lesson so far assumed a reward arrived from somewhere. For the setting

where these methods are now most used, nothing arrives. A model writes an answer,

and no part of the world reports a number.

A task with no score

What is available is weaker and much cheaper: a person shown two answers can say

which they prefer. That is a comparison, not a score, and the previous lessons

all need a score. The bridge is to assume a score exists and fit it.

FIG 1The chance that one response is preferred
the chance a rater prefers the first response to the second
the logistic curve, turning an unbounded gap into a probability
the gap between the two fitted scores, the only thing a comparison can inform
Note what this assumption buys and what it costs. It buys a single function that can score a response on its own, which comparisons cannot do. It costs the assumption that preferences are consistent with one underlying quantity, which human preferences are not, and the fitted score absorbs that inconsistency as noise.
FIG 2How a score gap turns into a preference rate
0.000.300.600.901.20-4.0-2.00.02.04.0gap between the two fitted scores
the logistic curve
A gap of zero means a coin flip, which is the right answer for two responses nobody can choose between. The flatness at the edges matters during fitting: once the model already predicts a comparison with near certainty, getting it more certain earns almost nothing, so the gradient concentrates on the pairs people actually disagreed about.

Only the gaps are determined, never the level: adding ten to every score changes

no preference. The fitted reward therefore has no absolute meaning, which is one

more reason the baseline of the fifth lesson is not optional here.

The lesson stops here

3 more paragraphs to go

You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.

The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.

See the planThe contents

This is the reading half

Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.

The contents

The rest of this course

  1. 01What a Policy Is Trying to Maximise
  2. 02The Value of Being Somewhere, and What It Must Agree Withopening only
  3. 03How to Differentiate Something You Can Only Sampleopening only
  4. 04An Unbiased Estimate That Is Almost Uselessopening only
  5. 05Judge an Action Against What Was Expected, Not Against Zeroopening only
  6. 06One Batch, Several Updates, and the Correction That Makes It Legalopening only
  7. 07Limit the Change in Behaviour, Not the Change in Weightsopening only
  8. 08A Reward Nobody Wrote, and the Tether That Keeps It Honestyou are here

Read alongside