Limit the Change in Behaviour, Not the Change in Weights
Last timeReusing Runs Collected by an Older Policy
A step size limits movement in parameters, which is the wrong quantity, and two repairs exist: a constraint on how much the policy itself changes, and a one line objective that imitates it.
The previous lesson ended with a budget and no way to spend it. Reuse is safe
while the policy stays near the one that collected the batch, and the only
instrument available for staying near is a step size, which measures the wrong
thing.
Parameters are the wrong ruler
A step of a fixed size in parameters can do almost nothing to the policy or it
can destroy it, and which one happens depends on where the policy currently is.
Where it is nearly uniform, large moves change little. Where it has become
confident, the same move can take a probability from 0.9 to 0.1 and the resulting
behaviour has nothing to do with what was learned. Worse, this is where training
spends most of its time, because a policy that is working is a confident one.
- an average over the states the old policy actually visited, taken from the batch
- how surprised the old policy would be by the new one's choices, action by action
- the budget for how much behaviour may change in one step, a small fixed number such as 0.01
- improve the reweighted objective as much as the budget allows, rather than by a fixed distance
- the smaller of the two terms, which is what makes the limit one sided
- the plain reweighted term from the previous lesson
- the same ratio with anything outside the band replaced by the nearest edge, the band being typically a fifth wide either way
| ratio | with advantage +2 | with advantage -2 | |
|---|---|---|---|
| ratio 0.6, well below th | 0.6 | 1.2 | -1.2 |
| ratio 1.0, unmoved | 1.0 | 2.0 | -2.0 |
| ratio 1.2, at the upper | 1.2 | 2.4 | -2.4 |
| ratio 2.0, well above th | 2.0 | 2.4 | -4.0 |
Why the smaller of the two
Replacing the ratio by its clipped value alone would be a mistake: a bad action
driven to a ratio of two would be penalised only as far as the edge of the band,
and the gradient there is zero, so nothing would pull it back. The minimum is what
makes the flat region appear on the side where flatness is safe.
Both methods are still local. Neither guarantees anything about a policy that has
moved far, and both inherit the state mismatch the previous lesson noted, which
no amount of correction on the actions can fix. What they buy is that a batch can
be used several times without the occasional catastrophic step that made the
plain estimator unusable.
The lesson stops here
1 more paragraph to go
You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.
The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.
See the planThe contentsThis is the reading half
Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.
The contents