Course Content
Generative AI System Design Interview
11 sections · 27 lessons
How image models generate: diffusion, GANs and VAEs
This is the most important explanation in the course's first half. Sections 8, 9, 10, and 11 are all diffusion case studies, and none of them makes sense without this one. Read it slowly.
The idea in one sentence: teach a network to remove a small amount of noise from an image, then generate a new image by starting from pure noise and applying that network many times.
The forward process — destroying an image on purpose
Start with a real photograph. Add a small amount of random noise. The result is the same photograph, slightly grainy. Add a small amount more. And again. After enough steps — a few hundred to a thousand, depending on the schedule — nothing of the original remains and you are looking at pure static.
Two things about this forward process matter enormously:
- It involves no learning. It is a fixed recipe: at step t, add this much noise. You can run it on any image at any time, for free.
- It gives you unlimited perfectly-labelled training pairs. For any image and any step, you know exactly what the noisy version looks like and exactly what noise you added. That is a supervised learning problem with a free, exact label — which is why diffusion training is so much more stable than the adversarial training described later in this lesson.
The training objective, stated plainly
Repeat millions of times:
- Pick a training image at random.
- Pick a step t at random, from 1 to T.
- Generate the noise for that step and add it to the image.
- Show the network the noisy image and the number t.
- Ask: what noise was added?
- Score the answer by how far the prediction is from the actual noise, and nudge the network's weights to reduce that gap.
That is the whole objective. It rewards one thing: accurately predicting the noise in a noisy image. There is no adversary, no discriminator, no game — one network, one regression target.
The reverse process — generating
Now generate. Start with an image of pure random noise, which contains no information at all.
- Ask the network what noise it sees.
- Remove a portion of the predicted noise.
- Add back a small amount of fresh random noise (most samplers do this; it keeps the process from collapsing onto an over-smoothed average).
- Repeat for N steps, with the amount removed and the amount added shrinking as you go.
After N steps you have an image. Which image? One determined by the random noise you started from — a different starting noise gives a different picture — plus any conditioning you supplied, which is what Section 9 (Text-to-Image Generation) adds.
Why the steps go coarse to fine
At high noise, almost nothing survives, so the only thing the network can predict is broad structure — where the dark mass is, roughly where the horizon sits. At low noise, the structure is already fixed and the only thing left to recover is fine detail. The step sequence therefore behaves like layout first, then shapes, then texture. Nobody programmed that ordering; it falls out of the noise schedule.
The analogy, and its three punctures
Diffusion is often described as sculpting a statue from marble — removing everything that is not the statue. Good opening. Now the punctures, because each one is a real misconception:
- The sculptor knows the statue in advance. Diffusion does not. There is no target image. The output is determined by the starting noise plus the conditioning, and both are inputs, not a hidden goal.
- Sculpting is pure removal. Diffusion is not. Most samplers add fresh noise back at each step. It is a stochastic walk, not a monotone subtraction.
- Nothing is "inside the marble". The same starting noise with a different text prompt produces a different image. The information comes from the model's weights and the conditioning, not from the noise.
Why this is slow, and what latent diffusion changes
One image needs N forward passes through a large network. At 25 steps and 40 ms per step, that is one second per image on a dedicated accelerator. At 50 steps it is two. Compare that with a classifier at 10 ms and you can see why cost dominates every image lesson here.
The fix that changed the field is latent diffusion: instead of running the process on pixels, first compress the image with an autoencoder into a much smaller array, run the entire diffusion process there, then decode once at the end. A 512 × 512 × 3 image is about 786,000 values; a typical compressed representation is around 64 × 64 × 4, or about 16,000 — roughly 48 times smaller. The per-step cost falls by a similar order. Latent diffusion, and why it changed everything covers what this costs in quality and where the losses show up.
GANs, VAEs, and why diffusion mostly displaced them
Diffusion is the default for image generation today, but it is not the only family, and Section 7 (Realistic Face Generation) is built on a generative adversarial network. You need all three to answer "why diffusion?" with anything better than "it is what everyone uses".
Generative adversarial networks
Two networks trained against each other.
- The generator takes a random vector and produces an image.
- The discriminator takes an image and predicts whether it came from the training set or from the generator.
The discriminator is trained to get better at telling them apart. The generator is trained to make images the discriminator misclassifies. Neither has an absolute target — each is chasing the other, which is what "adversarial" means.
The payoff is speed. Once trained, generation is one forward pass — tens of milliseconds, not the tens of steps diffusion needs. That is why GANs have not disappeared: they still win where latency is the binding constraint, such as real-time video effects and fast upscaling.
The cost is training. Two failure modes are characteristic:
- Mode collapse. The generator finds one output, or a narrow family of them, that fools the discriminator, and produces only that. You get beautiful faces that are all the same face. The model has stopped covering the data distribution and nothing in the loss says so.
- Instability. Because the two losses are defined against each other, neither number tells you whether the output is good. Losses can look healthy while the images are noise. You find out by looking, or by computing a distribution-level metric like Fréchet Inception Distance (see Metrics for images).
Variational autoencoders
A different bargain. An encoder compresses an image into a compact distribution in a latent space; a decoder reconstructs the image from a sample of it. Training rewards two things at once: reconstruct the input accurately, and keep the latent distribution close to a simple reference distribution so you can sample from it later.
Training is stable, the latent space is well-behaved and smooth, and samples are usually blurry. The blur is structural, not a tuning failure: when several outputs are plausible, a reconstruction objective is minimised by producing something near their average, and the average of several sharp images is a soft one.
Variational autoencoders did not lose. They moved. The compressor inside latent diffusion (covered above and in the latent diffusion lesson) is one — used for its stable, smooth compression, with the diffusion process supplying the sharpness it never had.
Why the field moved to diffusion
Four reasons, in rough order of importance:
- A stable objective. Predicting noise is ordinary supervised regression. There is no adversary to balance, and the loss curve means something.
- Distribution coverage. Diffusion is far less prone to mode collapse, so it produces genuinely varied output rather than variations on a favourite.
- It keeps improving with scale. More data and more compute reliably buy quality, which was much less dependable for adversarial training.
- Conditioning is easy. Injecting text, depth maps, or pose into the denoising network is a small architectural change (see Conditioning the diffusion process). Conditioning a GAN robustly on free text is harder.
And the price paid for all four: many forward passes per sample instead of one. The whole of Serving: making a slow model fast enough is about buying that back.
| GAN | VAE | Diffusion | |
|---|---|---|---|
| Inference | 1 pass — fastest | 1 pass | N passes — slowest |
| Training stability | Poor | Good | Good |
| Sample sharpness | High | Low (blurry) | High |
| Mode coverage | Weak | Good | Good |
| Text conditioning | Hard | Hard | Straightforward |
| Where it lives now | Real-time and upscaling; Section 7 | The compressor inside latent diffusion | The default for image, and the video in Section 11 |