What a Policy Is Trying to Maximise
A policy is a distribution over actions, a run of it earns a discounted total, and the quantity being maximised is the average of that total over every run the policy could produce.
Supervised learning is told the answer. Something is predicted, the correct value
is already written down beside it, and the gradient is what you get for being
wrong. Learning from reward has none of that. A sequence of decisions happens, a
number arrives, and nothing anywhere says which decision deserved it.
Before any of that can be fixed, the quantity being made large has to be written
down exactly, because almost every difficulty later is visible in it.
The loop
The policy is the only adjustable part. Writing it as a distribution rather than
as a choice is the first decision of the subject and it is not a modelling nicety.
If the policy always returned the single best action, a tiny change to the
parameters would either change nothing at all or flip the action outright, and a
quantity that moves in jumps has no useful derivative. Spreading probability
across the actions makes behaviour change smoothly with the parameters, which is
what makes everything after this possible.
- the reward collected at step t of this particular run
- the discount, a number between zero and one
- how many steps this run lasted
| step | step | reward | weight | contribution | what happened |
|---|---|---|---|---|---|
| 1 | 1 | 0 | 1 | 0 | Nothing yet, which is the usual case. Most steps in most tasks pay nothing. |
| 2 | 2 | 0 | 0.9 | 0 | Still nothing. The decision taken here may still be the one that mattered. |
| 3 | 3 | 1 | 0.81 | 0.81 | A reward arrives, two steps after the choice that set it up. |
| 4 | 4 | 2 | 0.729 | 1.458 | The largest reward counts for the least, because it came last. |
What the discount is doing
Two things at once, and it is worth keeping them apart.
The first is mathematical. Without it, a run that never ends has a sum that need
not converge, and comparing two infinite totals is not a well-posed question. The
discount makes the sum finite for any bounded reward.
The second is a modelling choice: it says how far ahead the learner is being
asked to care.
A smaller discount is also a noise control, which is the reason it is often set
well below what the task would justify. Shortening the horizon means each
decision is credited with fewer of the rewards that followed it, and fewer
rewards means less of the variance that the fourth lesson of this course is
entirely about. That is a bias traded for noise, deliberately.
- the parameters of the policy, the only thing being changed
- one whole run, states and actions together
- how likely that run is when this policy is the one acting
- the discounted return that run earned
The quantity being maximised
This is the sentence to hold on to, because almost everything that follows is a
consequence of it. In supervised learning the parameters sit inside the thing
being averaged, so the derivative passes straight through the average and lands
on the model. Here the parameters sit in the distribution being averaged over. You
cannot differentiate what you are sampling from by differentiating the samples.
The next lesson works out what a position is worth, which is the vocabulary the
repair needs, and the lesson after that performs the repair itself.
What to hold on to
A policy is a distribution over actions, kept random so that behaviour changes
smoothly with its parameters. A run earns one discounted number, and the discount
sets both whether the sum converges and how many steps ahead the learner is being
asked to care about. The objective is the average of that number over all the
runs the policy could produce, with the parameters sitting in the probabilities
rather than in the returns.
Recap
- A policy is a probability distribution over actions given a state, not a fixed plan, and its randomness is both where exploration comes from and what makes the objective differentiable at all.
- The return of a run is the discounted sum of the rewards collected along it, so it is a property of the whole run rather than of any single decision in it.
- The objective is the average return over all runs the policy could produce, which means the parameters being optimised sit inside the probability distribution rather than inside the quantity being averaged.
This is the reading half
Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.
The contents