The library

How models work

Derive how a single number arriving after a run of decisions can be turned into a gradient on every decision that produced it, what that gradient costs in noise, and what every method from the first estimator to the clipped objective behind modern fine-tuning is repairing

Learning From Reward

One number arrives at the end of a run and has to be shared out over every choice that led to it. This course derives the gradient that does it, the noise that makes it almost unusable, and each repair in turn.

8 lessons, written and corrected before you arrived. Reading them here needs no account. The first reads the whole way through; the others open and then stop, because a page nobody owns cannot tell who is reading it. Starting the course gives you your own copy, where every idea has problems standing under it and you can ask about any sentence.

Start reading

  1. 01What a Policy Is Trying to MaximiseA policy is a distribution over actions, a run of it earns a discounted total, and the quantity being maximised is the average of that total over every run the policy could produce.
  2. 02The Value of Being Somewhere, and What It Must Agree Withopening onlyThe value of a state is the average return from there onwards, the value of an action is the same thing after one forced choice, and both must agree with the values of whatever follows them.
  3. 03How to Differentiate Something You Can Only Sampleopening onlyOne identity turns the derivative of a distribution into an average under that distribution, which is what makes a gradient available from sampled runs without ever differentiating the environment.
  4. 04An Unbiased Estimate That Is Almost Uselessopening onlyThe sampled policy gradient is correct on average and still barely usable, because its noise grows with the length of the run, with the scale of the rewards and with how rarely anything is earned.
  5. 05Judge an Action Against What Was Expected, Not Against Zeroopening onlyAnything not depending on the action can be subtracted from the weight for free, and subtracting what the position was already worth turns a return into a judgement.
  6. 06One Batch, Several Updates, and the Correction That Makes It Legalopening onlyEvery estimator so far is thrown away after one step because the policy has changed, and a ratio of probabilities buys the right to reuse the batch until that ratio stops being trustworthy.
  7. 07Limit the Change in Behaviour, Not the Change in Weightsopening onlyA step size limits movement in parameters, which is the wrong quantity, and two repairs exist: a constraint on how much the policy itself changes, and a one line objective that imitates it.
  8. 08A Reward Nobody Wrote, and the Tether That Keeps It Honestopening onlyA language model has no score to collect, so the reward is fitted from human comparisons, and because a fitted reward can be gamed the objective carries a tether to the model it started from.