ContentsThe library

How Noise Becomes a Picture

Five Lines of Training, and One Squared Error

Last timeSkipping to Any Point on the Path

The entire training procedure is to pick a picture, pick a time, add the noise the closed form prescribes, and ask the network which noise it was.

Everything so far has been about the forward process, which involves no learning

at all. This lesson is the only place in the method where anything is trained,

and it is shorter than the lessons that set it up.

FIG 1What is minimised
the noise that was actually drawn and used, which is the answer
the network, given the noised picture and the time it is at
averaged over pictures, over times drawn uniformly along the path, and over draws of noise
This is the whole of training. No adversary, no second network, no sampling loop, and nothing that can collapse: it is an ordinary regression whose targets happen to be manufactured. The awkwardness in diffusion models is all in the sampler, and none of it is here.
FIG 2One training step
python
x0 = next(batch)                       # a batch of real pictures
t  = randint(1, T, size=len(x0))       # a different time for each one
e  = randn_like(x0)                    # the answer, drawn first

xt = sqrt(abar[t]) * x0 + sqrt(1 - abar[t]) * e
loss = mean((e - model(xt, t)) ** 2)
loss.backward()
Five lines, and the third one is the target. Note that the time differs across the batch, which is only possible because of the closed form from the previous lesson, and note that nothing in this loop ever runs the sampler. Training never produces a picture, and the quantity being minimised is not sample quality.
FIG 3A noise prediction is a picture prediction
the noised picture the network was given, which is known exactly
the network's guess at the noise
the implied guess at the clean picture, obtained by rearranging the closed form
The two targets carry identical information, so the choice between them is about numerical behaviour rather than about what is being learned. The noise target wins because it has the same scale at every point on the path, whereas the picture target is trivial near the start and hopeless near the end, which makes one loss impossible to balance across times.

One network for a thousand problems

The time is handed to the network as an input rather than being baked into

separate weights. That is not a saving of memory so much as a statement that the

problems at neighbouring times are nearly the same problem, so what is learned at

one of them transfers to its neighbours almost intact.

FIG 4How much each part of the path contributes
fraction along the pathshare of the true objectshare under the plain sq
almost clean0.0200.0010.090
early middle0.3000.1401.000
late middle0.7000.0850.610
almost pure noise0.9800.0040.030
The two right hand columns are the weighting that was dropped. The objective derived from first principles pours most of its weight into the two ends; the plain squared error spreads it across the middle instead. The marked cells are where the difference bites, and the method that produces better pictures is the one with no derivation behind it.

This is worth sitting with, because it is the first of several places in this

subject where the principled object and the one that works are not the same. The

honest description is that the simplified loss is a differently weighted version

of a bound, that the reweighting is not justified by any argument from first

principles, and that it was adopted because the pictures were better.

The lesson stops here

2 more paragraphs to go

You have read the opening. The rest of the argument, the problems that check whether it landed, and the lines worth keeping at the end all come with a plan.

The first lesson of every course in the library reads the whole way through, free, so you can see exactly what the rest of them are.

See the planThe contents

This is the reading half

Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.

The contents

The rest of this course

  1. 01A Thousand Small Steps From a Photograph to Static
  2. 02A Thousand Steps Collapse Into One Lineopening only
  3. 03Five Lines of Training, and One Squared Erroryou are here
  4. 04One Step Back, a Thousand Timesopening only
  5. 05The Network Is Pointing Uphillopening only
  6. 06Overshooting the Condition on Purposeopening only
  7. 07The Same Model, Twenty Times Fasteropening only
  8. 08If the Path Is Yours to Pick, Pick a Lineopening only

Read alongside