ContentsThe library

Learning From Reward

How to Differentiate Something You Can Only Sample

Last timeWhat a Position Is Worth

One identity turns the derivative of a distribution into an average under that distribution, which is what makes a gradient available from sampled runs without ever differentiating the environment.

The objective is an average over runs, with the parameters sitting in how likely

each run is rather than in what it earned. Differentiating the returns gives

exactly zero. One identity gets round this, and almost everything in the subject

is downstream of it.

FIG 1The log derivative identity
the probability of something under the current parameters
the gradient with respect to those parameters
the gradient of the log probability, often called the score
This is nothing more than the chain rule read backwards, since the gradient of the log is the gradient divided by the probability. Its value is structural: the right side has a probability multiplying something, which is the shape of an average, and the left side does not.

The identity

A gradient of a probability cannot be estimated by sampling, because samples tell

you what happened, not how likely it was to happen. The right side is an average

of the score under the very distribution being sampled from, and averages are

precisely what samples estimate. That is the whole move.

FIG 2The gradient of the objective
averaged over runs produced by the current policy, which is what sampling gives
how well this run went, a plain number, with no derivative taken of it anywhere
the direction that makes this entire run more likely
Read what this says before reading how it is computed. To improve, make runs that went well more likely and runs that went badly less likely. At no point is any action declared correct, which is the structural difference from supervised learning.

Where the environment goes

The probability of a whole run is the product of everything that had to happen:

each action chosen by the policy, and each state the environment produced in

reply. The logarithm turns that product into a sum.

The lesson stops here

4 more paragraphs to go

You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.

The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.

See the planThe contents

This is the reading half

Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.

The contents

The rest of this course

  1. 01What a Policy Is Trying to Maximise
  2. 02The Value of Being Somewhere, and What It Must Agree Withopening only
  3. 03How to Differentiate Something You Can Only Sampleyou are here
  4. 04An Unbiased Estimate That Is Almost Uselessopening only
  5. 05Judge an Action Against What Was Expected, Not Against Zeroopening only
  6. 06One Batch, Several Updates, and the Correction That Makes It Legalopening only
  7. 07Limit the Change in Behaviour, Not the Change in Weightsopening only
  8. 08A Reward Nobody Wrote, and the Tether That Keeps It Honestopening only

Read alongside