Generative AI System Design Interview

Course Content

Generative AI System Design Interview

11 sections · 27 lessons

How image models generate: diffusion, GANs and VAEs


This is the most important explanation in the course's first half. Sections 8, 9, 10, and 11 are all diffusion case studies, and none of them makes sense without this one. Read it slowly.

The idea in one sentence: teach a network to remove a small amount of noise from an image, then generate a new image by starting from pure noise and applying that network many times.

The forward process — destroying an image on purpose

Start with a real photograph. Add a small amount of random noise. The result is the same photograph, slightly grainy. Add a small amount more. And again. After enough steps — a few hundred to a thousand, depending on the schedule — nothing of the original remains and you are looking at pure static.

Two things about this forward process matter enormously:

  • It involves no learning. It is a fixed recipe: at step t, add this much noise. You can run it on any image at any time, for free.
  • It gives you unlimited perfectly-labelled training pairs. For any image and any step, you know exactly what the noisy version looks like and exactly what noise you added. That is a supervised learning problem with a free, exact label — which is why diffusion training is so much more stable than the adversarial training described later in this lesson.

The training objective, stated plainly

Repeat millions of times:

  1. Pick a training image at random.
  2. Pick a step t at random, from 1 to T.
  3. Generate the noise for that step and add it to the image.
  4. Show the network the noisy image and the number t.
  5. Ask: what noise was added?
  6. Score the answer by how far the prediction is from the actual noise, and nudge the network's weights to reduce that gap.

That is the whole objective. It rewards one thing: accurately predicting the noise in a noisy image. There is no adversary, no discriminator, no game — one network, one regression target.

The reverse process — generating

Now generate. Start with an image of pure random noise, which contains no information at all.

  1. Ask the network what noise it sees.
  2. Remove a portion of the predicted noise.
  3. Add back a small amount of fresh random noise (most samplers do this; it keeps the process from collapsing onto an over-smoothed average).
  4. Repeat for N steps, with the amount removed and the amount added shrinking as you go.

After N steps you have an image. Which image? One determined by the random noise you started from — a different starting noise gives a different picture — plus any conditioning you supplied, which is what Section 9 (Text-to-Image Generation) adds.

Forward: add noise on a schedule. Reverse: learn to remove it.x₀ — the image+ noisex₂₅₀+ noisex₅₀₀+ noisex₇₅₀+ noisex₁₀₀₀ — pure noiseFORWARD — fixed, no learningthe model predicts the noisethe model predicts the noisethe model predicts the noisethe model predicts the noiseREVERSE — this is the networkTraining is deliberately easy: take a real image, add a known amount of noise, and ask the network what noise was added. The label is free.Generation is then just the reverse walk from pure noise, repeated — which is why sampling steps are the dial that trades quality against cost.
Only the reverse arrows are learned — the forward process is a fixed schedule that manufactures training pairs for free.

Why the steps go coarse to fine

At high noise, almost nothing survives, so the only thing the network can predict is broad structure — where the dark mass is, roughly where the horizon sits. At low noise, the structure is already fixed and the only thing left to recover is fine detail. The step sequence therefore behaves like layout first, then shapes, then texture. Nobody programmed that ordering; it falls out of the noise schedule.

The analogy, and its three punctures

Diffusion is often described as sculpting a statue from marble — removing everything that is not the statue. Good opening. Now the punctures, because each one is a real misconception:

  • The sculptor knows the statue in advance. Diffusion does not. There is no target image. The output is determined by the starting noise plus the conditioning, and both are inputs, not a hidden goal.
  • Sculpting is pure removal. Diffusion is not. Most samplers add fresh noise back at each step. It is a stochastic walk, not a monotone subtraction.
  • Nothing is "inside the marble". The same starting noise with a different text prompt produces a different image. The information comes from the model's weights and the conditioning, not from the noise.

Why this is slow, and what latent diffusion changes

One image needs N forward passes through a large network. At 25 steps and 40 ms per step, that is one second per image on a dedicated accelerator. At 50 steps it is two. Compare that with a classifier at 10 ms and you can see why cost dominates every image lesson here.

The fix that changed the field is latent diffusion: instead of running the process on pixels, first compress the image with an autoencoder into a much smaller array, run the entire diffusion process there, then decode once at the end. A 512 × 512 × 3 image is about 786,000 values; a typical compressed representation is around 64 × 64 × 4, or about 16,000 — roughly 48 times smaller. The per-step cost falls by a similar order. Latent diffusion, and why it changed everything covers what this costs in quality and where the losses show up.

GANs, VAEs, and why diffusion mostly displaced them

Diffusion is the default for image generation today, but it is not the only family, and Section 7 (Realistic Face Generation) is built on a generative adversarial network. You need all three to answer "why diffusion?" with anything better than "it is what everyone uses".

What diffusion bought and what it costGAN and VAE• One forward pass, so generation is fast• Adversarial trainingcollapses or diverges• VAE samples come out blurryDiffusion• Thirty-plus passes,so generation is slow• A stable regression objective every step• Better coverage of the data distribution
Diffusion traded inference speed for training stability, and the field decided that was worth it.

Generative adversarial networks

Two networks trained against each other.

  • The generator takes a random vector and produces an image.
  • The discriminator takes an image and predicts whether it came from the training set or from the generator.

The discriminator is trained to get better at telling them apart. The generator is trained to make images the discriminator misclassifies. Neither has an absolute target — each is chasing the other, which is what "adversarial" means.

The payoff is speed. Once trained, generation is one forward pass — tens of milliseconds, not the tens of steps diffusion needs. That is why GANs have not disappeared: they still win where latency is the binding constraint, such as real-time video effects and fast upscaling.

The cost is training. Two failure modes are characteristic:

  • Mode collapse. The generator finds one output, or a narrow family of them, that fools the discriminator, and produces only that. You get beautiful faces that are all the same face. The model has stopped covering the data distribution and nothing in the loss says so.
  • Instability. Because the two losses are defined against each other, neither number tells you whether the output is good. Losses can look healthy while the images are noise. You find out by looking, or by computing a distribution-level metric like Fréchet Inception Distance (see Metrics for images).

Variational autoencoders

A different bargain. An encoder compresses an image into a compact distribution in a latent space; a decoder reconstructs the image from a sample of it. Training rewards two things at once: reconstruct the input accurately, and keep the latent distribution close to a simple reference distribution so you can sample from it later.

Training is stable, the latent space is well-behaved and smooth, and samples are usually blurry. The blur is structural, not a tuning failure: when several outputs are plausible, a reconstruction objective is minimised by producing something near their average, and the average of several sharp images is a soft one.

Variational autoencoders did not lose. They moved. The compressor inside latent diffusion (covered above and in the latent diffusion lesson) is one — used for its stable, smooth compression, with the diffusion process supplying the sharpness it never had.

Why the field moved to diffusion

Four reasons, in rough order of importance:

  1. A stable objective. Predicting noise is ordinary supervised regression. There is no adversary to balance, and the loss curve means something.
  2. Distribution coverage. Diffusion is far less prone to mode collapse, so it produces genuinely varied output rather than variations on a favourite.
  3. It keeps improving with scale. More data and more compute reliably buy quality, which was much less dependable for adversarial training.
  4. Conditioning is easy. Injecting text, depth maps, or pose into the denoising network is a small architectural change (see Conditioning the diffusion process). Conditioning a GAN robustly on free text is harder.

And the price paid for all four: many forward passes per sample instead of one. The whole of Serving: making a slow model fast enough is about buying that back.

GANVAEDiffusion
Inference1 pass — fastest1 passN passes — slowest
Training stabilityPoorGoodGood
Sample sharpnessHighLow (blurry)High
Mode coverageWeakGoodGood
Text conditioningHardHardStraightforward
Where it lives nowReal-time and upscaling; Section 7The compressor inside latent diffusionThe default for image, and the video in Section 11