Derive how a single number arriving after a run of decisions can be turned into a gradient on every decision that produced it, what that gradient costs in noise, and what every method from the first estimator to the clipped objective behind modern fine-tuning is repairing
Learning From Reward

One number arrives at the end of a run and has to be shared out over every choice that led to it. This course derives the gradient that does it, the noise that makes it almost unusable, and each repair in turn.
8 lessons, written and corrected before you arrived. Reading them here needs no account. The first reads the whole way through; the others open and then stop, because a page nobody owns cannot tell who is reading it. Starting the course gives you your own copy, where every idea has problems standing under it and you can ask about any sentence.