ContentsThe library

Learning From Reward

An Unbiased Estimate That Is Almost Useless

Last timeA Gradient Through a Sample

The sampled policy gradient is correct on average and still barely usable, because its noise grows with the length of the run, with the scale of the rewards and with how rarely anything is earned.

There is now a gradient that can be computed from sampled runs. It is correct on

average. It is also, used directly, close to unusable, and understanding exactly

why is what makes every later method look inevitable rather than clever.

FIG 1The estimate formed from a batch
how many runs were collected before an update
what the ith run earned after step t, the weight on that action
the direction that makes that action more likely
This average is the only contact anything has with the true gradient. Every property of training, good or bad, is a property of this quantity, which is why it is worth examining rather than implementing.

Unbiased, and what that does not mean

Unbiased means that if this were computed infinitely many times and averaged, the

result would be the true gradient. Nothing in that sentence says anything about

what happens with thirty-two runs.

FIG 2Three runs of the same policy on the same task
steprunreturnweight on every action in that runwhat happened
1100Nothing earned. Every action in this run is pushed by zero, so the run taught nothing at all.
2200Again nothing. Two thirds of the batch has contributed nothing to the update.
3311One run found something. The entire update is now whatever this single run happened to do, including everything irrelevant it did along the way.
3 steps
Sparse reward is the ordinary case rather than a hard case. When almost every run earns nothing, the estimate is formed from the few that did, and every incidental action in those runs is reinforced alongside the ones that mattered.

Where the spread comes from

Three independent sources of randomness enter the single number each action is

weighted by. The action draws themselves, since the policy is a distribution. The

environment's answers. And the rewards, when they are noisy. These do not cancel;

they accumulate along the run, so the spread of the weight grows with how long the

run is, even when the task is behaving perfectly.

The lesson stops here

4 more paragraphs to go

You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.

The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.

See the planThe contents

This is the reading half

Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.

The contents

The rest of this course

  1. 01What a Policy Is Trying to Maximise
  2. 02The Value of Being Somewhere, and What It Must Agree Withopening only
  3. 03How to Differentiate Something You Can Only Sampleopening only
  4. 04An Unbiased Estimate That Is Almost Uselessyou are here
  5. 05Judge an Action Against What Was Expected, Not Against Zeroopening only
  6. 06One Batch, Several Updates, and the Correction That Makes It Legalopening only
  7. 07Limit the Change in Behaviour, Not the Change in Weightsopening only
  8. 08A Reward Nobody Wrote, and the Tether That Keeps It Honestopening only

Read alongside