ContentsThe library

How Noise Becomes a Picture

The Network Is Pointing Uphill

Last timeWalking the Path Backwards

The predicted noise turns out to be a scaled gradient of the log density of noised pictures, which is why a denoiser can be used as a generator at all.

So far the trained network has been described by what it was asked to do: name

the noise. This lesson works out what that answer is, in terms that have nothing

to do with noise, and the answer explains why a loop of denoising steps produces

pictures rather than merely cleaning them.

A density for every noise level

Each time on the path carries its own distribution: the distribution of real

pictures with that much noise added. At the start of the path this is almost the

distribution of real pictures. At the end it is almost pure noise, the same for

every dataset. In between is a sequence of progressively blurred versions of one

distribution.

FIG 1One distribution, blurred by the amount on the dial
00.250.50.7511.251.5-6-4-20246a single direction through the space of pictures
real pictures, blurred by s
Two kinds of picture, sitting at plus and minus two. At the smallest blur there are two needles with nothing between or beyond them, and a point out at five has no slope to tell it which needle to head for. Turn the dial up and the two merge into one broad mound whose slope reaches everywhere, but which no longer knows there were two kinds. The sampler needs both ends of this dial, which is why it walks down the path rather than picking one noise level.
FIG 2The gradient of the log density, and what the network computes
the distribution of real pictures blurred by the noise belonging to time t
the direction from this point that most increases plausibility, with length saying how steeply
the trained prediction of the noise present in this input
the noise level at this time, which is the only thing converting between the two
The equality is exact for a perfectly trained network, and it is a statement about units rather than an approximation. Predicting the noise and computing the uphill direction are the same calculation; the minus sign is just the observation that removing the noise and increasing plausibility point the same way.

What the gradient of a log density is

Two features of this quantity are worth stating separately, because they are the

reason it is the useful object rather than the density itself. First, it survives

differentiation of the logarithm without any normalising constant: multiply the

density by any constant and the gradient of its logarithm is unchanged. A

quantity that does not need the constant can be estimated from samples, and the

constant is exactly the part that is impossible to compute. Second, it is a

direction at every point in the space, which is a thing a loop can follow.

The lesson stops here

3 more paragraphs to go

You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.

The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.

See the planThe contents

This is the reading half

Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.

The contents

The rest of this course

  1. 01A Thousand Small Steps From a Photograph to Static
  2. 02A Thousand Steps Collapse Into One Lineopening only
  3. 03Five Lines of Training, and One Squared Erroropening only
  4. 04One Step Back, a Thousand Timesopening only
  5. 05The Network Is Pointing Uphillyou are here
  6. 06Overshooting the Condition on Purposeopening only
  7. 07The Same Model, Twenty Times Fasteropening only
  8. 08If the Path Is Yours to Pick, Pick a Lineopening only

Read alongside