ContentsThe library

Learning From Reward

What a Policy Is Trying to Maximise

A policy is a distribution over actions, a run of it earns a discounted total, and the quantity being maximised is the average of that total over every run the policy could produce.

Supervised learning is told the answer. Something is predicted, the correct value

is already written down beside it, and the gradient is what you get for being

wrong. Learning from reward has none of that. A sequence of decisions happens, a

number arrives, and nothing anywhere says which decision deserved it.

Before any of that can be fixed, the quantity being made large has to be written

down exactly, because almost every difficulty later is visible in it.

FIG 1One step of the interaction
The policy owns one box in this picture. It cannot change how the environment answers, which is exactly why the gradient worked out in a later lesson never needs to differentiate through it.

The loop

The policy is the only adjustable part. Writing it as a distribution rather than

as a choice is the first decision of the subject and it is not a modelling nicety.

If the policy always returned the single best action, a tiny change to the

parameters would either change nothing at all or flip the action outright, and a

quantity that moves in jumps has no useful derivative. Spreading probability

across the actions makes behaviour change smoothly with the parameters, which is

what makes everything after this possible.

FIG 2The discounted return of one run
the reward collected at step t of this particular run
the discount, a number between zero and one
how many steps this run lasted
One number for one run. Note what is absent: any mention of which action was good. The return is earned by the run as a whole, and taking it apart is the problem the rest of this course solves.
FIG 3A four step run, discounted at 0.9
stepsteprewardweightcontributionwhat happened
11010Nothing yet, which is the usual case. Most steps in most tasks pay nothing.
2200.90Still nothing. The decision taken here may still be the one that mattered.
3310.810.81A reward arrives, two steps after the choice that set it up.
4420.7291.458The largest reward counts for the least, because it came last.
4 steps
The return of this run is 2.268. Every step in it shares that number equally as far as the arithmetic is concerned, and the first two steps paid nothing at all, which is why a method that only looks at the rewards next to a decision learns nothing here.

What the discount is doing

Two things at once, and it is worth keeping them apart.

The first is mathematical. Without it, a run that never ends has a sum that need

not converge, and comparing two infinite totals is not a well-posed question. The

discount makes the sum finite for any bounded reward.

The second is a modelling choice: it says how far ahead the learner is being

asked to care.

FIG 4How far ahead the discount looks
0.0062.50125.00187.50250.000.50.60.70.91.0discount
one over one minus the discount
The horizon is not linear in the discount, which is the practical point. Going from 0.9 to 0.99 does not extend the horizon by a tenth; it multiplies it by ten, and a run of a hundred steps that was being credited over ten is suddenly being credited over all of it.

A smaller discount is also a noise control, which is the reason it is often set

well below what the task would justify. Shortening the horizon means each

decision is credited with fewer of the rewards that followed it, and fewer

rewards means less of the variance that the fourth lesson of this course is

entirely about. That is a bias traded for noise, deliberately.

FIG 5The objective
the parameters of the policy, the only thing being changed
one whole run, states and actions together
how likely that run is when this policy is the one acting
the discounted return that run earned
Read where the parameters are. They are not in the return, which is a fact about rewards the environment handed out. They are in the probability of the run happening at all, and that is the structural difficulty of the whole subject.

The quantity being maximised

This is the sentence to hold on to, because almost everything that follows is a

consequence of it. In supervised learning the parameters sit inside the thing

being averaged, so the derivative passes straight through the average and lands

on the model. Here the parameters sit in the distribution being averaged over. You

cannot differentiate what you are sampling from by differentiating the samples.

The next lesson works out what a position is worth, which is the vocabulary the

repair needs, and the lesson after that performs the repair itself.

What to hold on to

A policy is a distribution over actions, kept random so that behaviour changes

smoothly with its parameters. A run earns one discounted number, and the discount

sets both whether the sum converges and how many steps ahead the learner is being

asked to care about. The objective is the average of that number over all the

runs the policy could produce, with the parameters sitting in the probabilities

rather than in the returns.

Recap

  • A policy is a probability distribution over actions given a state, not a fixed plan, and its randomness is both where exploration comes from and what makes the objective differentiable at all.
  • The return of a run is the discounted sum of the rewards collected along it, so it is a property of the whole run rather than of any single decision in it.
  • The objective is the average return over all runs the policy could produce, which means the parameters being optimised sit inside the probability distribution rather than inside the quantity being averaged.

This is the reading half

Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.

The contents

NextWhat a Position Is Worth →

The rest of this course

  1. 01What a Policy Is Trying to Maximiseyou are here
  2. 02The Value of Being Somewhere, and What It Must Agree Withopening only
  3. 03How to Differentiate Something You Can Only Sampleopening only
  4. 04An Unbiased Estimate That Is Almost Uselessopening only
  5. 05Judge an Action Against What Was Expected, Not Against Zeroopening only
  6. 06One Batch, Several Updates, and the Correction That Makes It Legalopening only
  7. 07Limit the Change in Behaviour, Not the Change in Weightsopening only
  8. 08A Reward Nobody Wrote, and the Tether That Keeps It Honestopening only

Read alongside