A Thousand Small Steps From a Photograph to Static
The forward process shrinks the picture slightly and adds a little noise, a thousand times over, and three properties of that one line are what make the whole method possible.
Generating a picture from nothing is hard. Destroying one is easy, and more
usefully, destroying one is exactly describable. The method in this course is
built on the gap between those two facts.
- the picture as it stood a moment ago, which may already be partly destroyed
- how much noise this particular step adds, a small fixed number such as 0.0001 near the start
- fresh noise, drawn independently for this step and never reused
Why the shrink
The shrink is the part that looks like an extra detail and is not. Without it,
each step would add variance to what was already there and the values would grow
without limit, so where the path ended would depend on how long it was run.
| adding noise only | shrinking first | |
|---|---|---|
| at the start | 1.00 | 1.00 |
| after 1 step | 1.02 | 1.00 |
| after 10 steps | 1.20 | 1.00 |
| after 100 steps | 3.00 | 1.00 |
| step | step | share of the original left | share that is noise | what happened |
|---|---|---|---|---|
| 1 | 0 | 1 | 0 | The picture itself. |
| 2 | 50 | 0.98 | 0.02 | Visually identical. The model learns almost nothing here, which is a real cost of a badly chosen schedule. |
| 3 | 300 | 0.7 | 0.3 | Texture is gone, shapes and colours remain. This is where most of what matters is learned. |
| 4 | 700 | 0.22 | 0.78 | Only the broadest arrangement survives. Early sampling steps come from here. |
| 5 | 1000 | 0.01 | 0.99 | Nothing. The same for every picture, which is exactly what is needed. |
No memory
Each step depends only on the step immediately before it. That is worth stating
plainly because both of the next two lessons are consequences of it. The forward
chain collapses into a single formula connecting the original picture to any
point on the path, which is what makes training affordable. And the reverse
becomes a sequence of small local guesses rather than one global inversion, which
is what makes it learnable at all.
What to hold on to
The forward process shrinks and adds noise, one small step at a time, with no
network and nothing learned. The shrink is what keeps the spread from growing, so
the path has a settled destination rather than a moving one. The amount added per
step is a fixed schedule, and choosing it decides where along the path the real
structure is destroyed. The destination is standard noise that has forgotten
which picture it came from, which is the only thing a sampler can be asked to
start from.
Recap
- Each forward step shrinks what is there and adds fresh noise, and the shrink is chosen so that the total spread of the values stays put rather than growing without limit.
- The amount of noise added per step is a schedule, chosen in advance and never learned, and it decides at which stage of the path the picture's structure is actually destroyed.
- After enough steps the result is standard noise with nothing of the original left, which is the property that lets sampling begin from noise that nobody had to generate from a picture.
This is the reading half
Starting the course gives you your own copy of it. Every idea on every page has problems standing under it, marked with a reason rather than a tick, and any sentence you do not believe can be opened and argued with. None of that can happen on a page nobody owns.
The contents