ContentsThe library

Learning From Reward

Limit the Change in Behaviour, Not the Change in Weights

Last timeReusing Runs Collected by an Older Policy

A step size limits movement in parameters, which is the wrong quantity, and two repairs exist: a constraint on how much the policy itself changes, and a one line objective that imitates it.

The previous lesson ended with a budget and no way to spend it. Reuse is safe

while the policy stays near the one that collected the batch, and the only

instrument available for staying near is a step size, which measures the wrong

thing.

Parameters are the wrong ruler

A step of a fixed size in parameters can do almost nothing to the policy or it

can destroy it, and which one happens depends on where the policy currently is.

Where it is nearly uniform, large moves change little. Where it has become

confident, the same move can take a probability from 0.9 to 0.1 and the resulting

behaviour has nothing to do with what was learned. Worse, this is where training

spends most of its time, because a policy that is working is a confident one.

FIG 1How far the policy has moved
an average over the states the old policy actually visited, taken from the batch
how surprised the old policy would be by the new one's choices, action by action
Zero when the two policies agree everywhere, and growing as they disagree. Crucially it does not care how the parameters are arranged: two networks that behave identically are at distance zero from one another however different their weights. That is the property a step size cannot have.
FIG 2The step as a constrained problem
the budget for how much behaviour may change in one step, a small fixed number such as 0.01
improve the reweighted objective as much as the budget allows, rather than by a fixed distance
This is the whole method in one line: take the largest improvement available without changing behaviour by more than the budget. Its solution is the ordinary gradient rescaled by the curvature of the divergence, which is a matrix the size of the parameter count squared, approached by iterative solvers rather than formed. Correct, and heavy.
FIG 3The clipped objective
the smaller of the two terms, which is what makes the limit one sided
the plain reweighted term from the previous lesson
the same ratio with anything outside the band replaced by the nearest edge, the band being typically a fifth wide either way
There is no constraint here and no second order object. The policy is free to move anywhere; the objective simply stops increasing once the ratio leaves the band, so an ordinary optimiser has no reason to take it further. That is a weaker promise than the constraint, in exchange for being a single line of arithmetic.
FIG 4Where the band sits, for a clip width of 0.2
a good action pulls as normal, a bad one pays nothing furtherthe objective behaves exactly as the unclipped onea good action pays nothing further, a bad one pushes as normal0.50.751.251.51.751the policy that collected the batch0.8lower edge1.2upper edgeratio of new probability to old, for one action
Note that the two outer regions are not mirror images in effect. Which side goes flat depends on the sign of the advantage, which is the whole point of taking a minimum rather than simply replacing the ratio by its clipped value.
FIG 5The clipped term at four ratios, for a good and a bad action
ratiowith advantage +2with advantage -2
ratio 0.6, well below th0.61.2-1.2
ratio 1.0, unmoved1.02.0-2.0
ratio 1.2, at the upper 1.22.4-2.4
ratio 2.0, well above th2.02.4-4.0
The last row is the asymmetry. Pushed far past the upper edge, a good action has stopped paying anything more, so the optimiser has no reason to continue. A bad action at the same place is still being penalised at full strength, and the penalty grows, so a step that drove a bad action up can always be walked back. The limit applies to enthusiasm and not to correction.

Why the smaller of the two

Replacing the ratio by its clipped value alone would be a mistake: a bad action

driven to a ratio of two would be penalised only as far as the edge of the band,

and the gradient there is zero, so nothing would pull it back. The minimum is what

makes the flat region appear on the side where flatness is safe.

Both methods are still local. Neither guarantees anything about a policy that has

moved far, and both inherit the state mismatch the previous lesson noted, which

no amount of correction on the actions can fix. What they buy is that a batch can

be used several times without the occasional catastrophic step that made the

plain estimator unusable.

The lesson stops here

1 more paragraph to go

You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.

The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.

See the planThe contents

This is the reading half

Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.

The contents

The rest of this course

  1. 01What a Policy Is Trying to Maximise
  2. 02The Value of Being Somewhere, and What It Must Agree Withopening only
  3. 03How to Differentiate Something You Can Only Sampleopening only
  4. 04An Unbiased Estimate That Is Almost Uselessopening only
  5. 05Judge an Action Against What Was Expected, Not Against Zeroopening only
  6. 06One Batch, Several Updates, and the Correction That Makes It Legalopening only
  7. 07Limit the Change in Behaviour, Not the Change in Weightsyou are here
  8. 08A Reward Nobody Wrote, and the Tether That Keeps It Honestopening only

Read alongside