ContentsThe library

Learning From Reward

Judge an Action Against What Was Expected, Not Against Zero

Last timeWhy the Estimate Is So Noisy

Anything not depending on the action can be subtracted from the weight for free, and subtracting what the position was already worth turns a return into a judgement.

The last lesson ended with a diagnosis. The weight on every action contains a

large part shared by all of them, carrying no information and all of the noise.

This lesson removes it, and the removal turns out to be free.

FIG 1The average score is zero
the probabilities of the available actions, which add to one by definition
the gradient of a constant, which is zero no matter how the parameters move
Everything in this lesson rests on this line. The policy's probabilities sum to one for every setting of the parameters, so the average direction that would make a drawn action likelier is nothing at all. Pushing all actions up in proportion to how likely they already are is not a change of policy.

Why anything can be subtracted

Multiply that zero average by any quantity that does not depend on which action

was drawn and it is still zero. So that quantity can be subtracted from every

weight without changing the gradient in the slightest.

FIG 2The estimator with a baseline
any number that depends on the state but not on the action taken there
what was earned, compared against what was expected rather than against zero
Note carefully what the baseline may depend on. The state, yes, and the time step, and anything else already settled before the action. The action itself, no: a baseline that looks at the action would cancel part of the signal rather than part of the noise, and the gradient would no longer be the right one.

What to subtract

The obvious candidate is what the position was already worth, which the second

lesson called the value of the state. Subtracting it leaves the part of the weight

that differs between the actions available at that moment, which is the only part

that says anything about the decision.

The lesson stops here

4 more paragraphs to go

You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.

The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.

See the planThe contents

This is the reading half

Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.

The contents

The rest of this course

  1. 01What a Policy Is Trying to Maximise
  2. 02The Value of Being Somewhere, and What It Must Agree Withopening only
  3. 03How to Differentiate Something You Can Only Sampleopening only
  4. 04An Unbiased Estimate That Is Almost Uselessopening only
  5. 05Judge an Action Against What Was Expected, Not Against Zeroyou are here
  6. 06One Batch, Several Updates, and the Correction That Makes It Legalopening only
  7. 07Limit the Change in Behaviour, Not the Change in Weightsopening only
  8. 08A Reward Nobody Wrote, and the Tether That Keeps It Honestopening only

Read alongside