ContentsThe library

How Noise Becomes a Picture

A Thousand Small Steps From a Photograph to Static

The forward process shrinks the picture slightly and adds a little noise, a thousand times over, and three properties of that one line are what make the whole method possible.

Generating a picture from nothing is hard. Destroying one is easy, and more

usefully, destroying one is exactly describable. The method in this course is

built on the gap between those two facts.

FIG 1A single forward step
the picture as it stood a moment ago, which may already be partly destroyed
how much noise this particular step adds, a small fixed number such as 0.0001 near the start
fresh noise, drawn independently for this step and never reused
There is no network here and nothing to learn. Every value in the picture is shrunk by the same factor and has its own fresh noise added. The entire forward process is this line applied a thousand times with a different number each time.
FIG 2The chain, and the two directions along it
The solid direction is free and exact. The dashed direction is the method. Note that the reverse is a chain of small steps rather than one leap, which is what keeps each individual guess easy enough to be learnable.

Why the shrink

The shrink is the part that looks like an extra detail and is not. Without it,

each step would add variance to what was already there and the values would grow

without limit, so where the path ended would depend on how long it was run.

FIG 3Spread of the values after repeated steps, with and without the shrink
adding noise onlyshrinking first
at the start1.001.00
after 1 step1.021.00
after 10 steps1.201.00
after 100 steps3.001.00
With the shrink, a spread of one is a fixed point: shrinking by the square root of one minus beta and adding beta of fresh variance returns exactly one. Without it the process never settles, and the marked cell is already somewhere no sampler could start from, since nobody knows which value it would have reached.
FIG 4Two schedules over a thousand steps
0.000.010.010.020.030.0250.0500.0750.01000.0step along the path
a straight line from 0.0001 to 0.02a schedule that waits longer before destroying
The schedule is fixed before training and never adjusted by it. Both of these end at the same place, and they differ in where the picture's structure is actually lost: under the straight line a surprising amount is already gone a third of the way along, which means much of the path is spent on noise that teaches the model very little.
FIG 5What survives, under the straight line schedule
stepstepshare of the original leftshare that is noisewhat happened
1010The picture itself.
2500.980.02Visually identical. The model learns almost nothing here, which is a real cost of a badly chosen schedule.
33000.70.3Texture is gone, shapes and colours remain. This is where most of what matters is learned.
47000.220.78Only the broadest arrangement survives. Early sampling steps come from here.
510000.010.99Nothing. The same for every picture, which is exactly what is needed.
5 steps
Reading this from the bottom up is reading the sampler's job. It starts at the last row with nothing, and each reverse step has to add back a little of the structure the corresponding forward step removed.

No memory

Each step depends only on the step immediately before it. That is worth stating

plainly because both of the next two lessons are consequences of it. The forward

chain collapses into a single formula connecting the original picture to any

point on the path, which is what makes training affordable. And the reverse

becomes a sequence of small local guesses rather than one global inversion, which

is what makes it learnable at all.

What to hold on to

The forward process shrinks and adds noise, one small step at a time, with no

network and nothing learned. The shrink is what keeps the spread from growing, so

the path has a settled destination rather than a moving one. The amount added per

step is a fixed schedule, and choosing it decides where along the path the real

structure is destroyed. The destination is standard noise that has forgotten

which picture it came from, which is the only thing a sampler can be asked to

start from.

Recap

  • Each forward step shrinks what is there and adds fresh noise, and the shrink is chosen so that the total spread of the values stays put rather than growing without limit.
  • The amount of noise added per step is a schedule, chosen in advance and never learned, and it decides at which stage of the path the picture's structure is actually destroyed.
  • After enough steps the result is standard noise with nothing of the original left, which is the property that lets sampling begin from noise that nobody had to generate from a picture.

This is the reading half

Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.

The contents

NextSkipping to Any Point on the Path →

The rest of this course

  1. 01A Thousand Small Steps From a Photograph to Staticyou are here
  2. 02A Thousand Steps Collapse Into One Lineopening only
  3. 03Five Lines of Training, and One Squared Erroropening only
  4. 04One Step Back, a Thousand Timesopening only
  5. 05The Network Is Pointing Uphillopening only
  6. 06Overshooting the Condition on Purposeopening only
  7. 07The Same Model, Twenty Times Fasteropening only
  8. 08If the Path Is Yours to Pick, Pick a Lineopening only

Read alongside