ContentsThe library

Learning From Reward

One Batch, Several Updates, and the Correction That Makes It Legal

Last timeSubtracting What Was Already Expected

Every estimator so far is thrown away after one step because the policy has changed, and a ratio of probabilities buys the right to reuse the batch until that ratio stops being trustworthy.

There is an awkward fact buried in every estimator so far. The gradient is an

average over runs that the current policy would produce. One step of training

later, the current policy is a different policy, and the expensively collected

batch is a set of samples from a distribution nobody cares about any more.

Strictly, one batch buys one step. Since collecting the batch means actually

running the environment, and taking the step means a single pass of arithmetic,

that is a terrible exchange rate.

FIG 1An average under one distribution from samples of another
the distribution whose average is wanted, here the policy being improved
the distribution the samples actually came from, here the policy that collected the batch
how much more often the wanted distribution would have produced this sample
Read from left to right it looks like a trick. Read from right to left it is bookkeeping: a sample that the wanted distribution would have produced twice as often is counted twice. The identity is exact, and it needs only that the sampling distribution gives some chance to everything the wanted one does.

The correction

Applied to a policy, the ratio is the new probability of the action that was

taken over the probability it had when it was taken. The second number is

already known, because the policy recorded it at collection time.

FIG 2The reweighted objective
an average over every step in the batch, not over fresh runs
what the policy being trained now says about the action that was taken
what it said at the moment the action was taken, recorded in the batch
the advantage of that action, computed once from the batch
Everything here except the first factor is a fixed number stored in the batch. That is what makes several passes possible: each pass recomputes one probability per step and nothing else.

Why its gradient is the right one

At the start of a pass the policy being trained is the policy that collected the

data, so every ratio is exactly one.

The lesson stops here

3 more paragraphs to go

You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.

The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.

See the planThe contents

This is the reading half

Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.

The contents

The rest of this course

  1. 01What a Policy Is Trying to Maximise
  2. 02The Value of Being Somewhere, and What It Must Agree Withopening only
  3. 03How to Differentiate Something You Can Only Sampleopening only
  4. 04An Unbiased Estimate That Is Almost Uselessopening only
  5. 05Judge an Action Against What Was Expected, Not Against Zeroopening only
  6. 06One Batch, Several Updates, and the Correction That Makes It Legalyou are here
  7. 07Limit the Change in Behaviour, Not the Change in Weightsopening only
  8. 08A Reward Nobody Wrote, and the Tether That Keeps It Honestopening only

Read alongside