ContentsThe library

Learning From Reward

The Value of Being Somewhere, and What It Must Agree With

Last timeThe Number Being Made Large

The value of a state is the average return from there onwards, the value of an action is the same thing after one forced choice, and both must agree with the values of whatever follows them.

The last lesson left one number for a whole run and no way to say which decision

earned it. Before that can be fixed, there has to be a way of saying what a

position is worth, so that a decision can be judged against what was already

expected rather than against nothing.

FIG 1What a position is worth
the value of state s under the policy pi
the discounted return collected from that point onwards
the policy that keeps acting for the rest of the run
Two things are fixed here and both matter. The state is where the run is, and the policy is what will happen next. Change either and the number changes, which is why a value function belongs to a policy and is invalidated the moment the policy moves.

The value of a state

A position is worth what tends to follow from it. Not what did follow once, which

is a sample of this quantity and can be wildly off it, but the average over

everything that could follow.

FIG 2What a position is worth after one forced choice
the value of taking action a in state s and continuing with pi
the first action is forced to be a, whatever the policy would have preferred
This is the quantity a learner actually wants. It holds the position fixed and varies only the decision, which is the comparison that says whether a choice was better or worse than the alternatives available at the same moment.

The value of a state and one decision

The relationship between the two is immediate: the value of a state is the

average of the action values over whatever the policy would have done there. If

every action in a state is worth the same, the policy has nothing to learn in

that state, however large the numbers are.

The lesson stops here

4 more paragraphs to go

You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.

The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.

See the planThe contents

This is the reading half

Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.

The contents

The rest of this course

  1. 01What a Policy Is Trying to Maximise
  2. 02The Value of Being Somewhere, and What It Must Agree Withyou are here
  3. 03How to Differentiate Something You Can Only Sampleopening only
  4. 04An Unbiased Estimate That Is Almost Uselessopening only
  5. 05Judge an Action Against What Was Expected, Not Against Zeroopening only
  6. 06One Batch, Several Updates, and the Correction That Makes It Legalopening only
  7. 07Limit the Change in Behaviour, Not the Change in Weightsopening only
  8. 08A Reward Nobody Wrote, and the Tether That Keeps It Honestopening only

Read alongside