A Reward Nobody Wrote, and the Tether That Keeps It Honest
Last timeKeeping the Step Small Where It Matters
A language model has no score to collect, so the reward is fitted from human comparisons, and because a fitted reward can be gamed the objective carries a tether to the model it started from.
Every lesson so far assumed a reward arrived from somewhere. For the setting
where these methods are now most used, nothing arrives. A model writes an answer,
and no part of the world reports a number.
A task with no score
What is available is weaker and much cheaper: a person shown two answers can say
which they prefer. That is a comparison, not a score, and the previous lessons
all need a score. The bridge is to assume a score exists and fit it.
- the chance a rater prefers the first response to the second
- the logistic curve, turning an unbounded gap into a probability
- the gap between the two fitted scores, the only thing a comparison can inform
Only the gaps are determined, never the level: adding ten to every score changes
no preference. The fitted reward therefore has no absolute meaning, which is one
more reason the baseline of the fifth lesson is not optional here.
The lesson stops here
3 more paragraphs to go
You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.
The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.
See the planThe contentsThis is the reading half
Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.
The contents