Course Content
Introduction to Generative AI
3 sections · 9 lessons
Diffusion Models
Take a photograph and add a small amount of random noise — enough to make it look slightly grainy. Now train a network to remove that grain. This is a solved problem, and has been for decades. It is ordinary regression: input a grainy image, output a clean one, minimise squared error. It trains in an afternoon and it works.
Now ask a bolder question. If a network can remove a little noise, what happens if you run it repeatedly? Start from an image that is 5% noise, denoise, and you have a clean image. Start from one that is 50% noise, denoise, denoise again, and again — does something emerge?
Push the idea to its limit. Start from pure noise, containing no image whatsoever, and apply the denoiser hundreds of times. Would a photograph appear out of static?
Remarkably, yes — but only if you fix one thing first. A denoiser trained on slightly grainy photographs has never seen pure static. Hand it a screen of random values and it produces nonsense, because that input is nothing like its training data. The repair is straightforward and it is the whole design: train one network across every noise level, from imperceptible grain to total static, and tell it at each call how noisy the input is supposed to be. Now it knows which job it is doing. At high noise it hallucinates coarse structure; at low noise it sharpens fine detail.
That is a diffusion model. Everything else is detail about how to add the noise, how to tell the network where it is, and how to take fewer steps.
The impossible leap from noise to a photograph becomes possible when it is broken into a thousand steps, each of which is a mild regression problem.
The forward process: destroying an image on purpose
Define a fixed recipe for corruption. At each step t, shrink the current image slightly and add a little Gaussian noise:
The βt values are small — typically from 10−4 up to 0.02 over T=1000 steps — and are chosen in advance, not learned. Run this to completion and any image becomes indistinguishable from pure noise.
Simulating a thousand steps for every training example would be crippling, so the key practical fact is that this process has a closed form. Writing αt=1−βt and αˉt=∏s=1tαs:
One line, one noise draw, and you jump directly to any step. αˉt is the fraction of the original signal that survives; 1−αˉt is the noise fraction. They sum to one, which is why the image's overall scale stays stable rather than exploding.
Concrete values for the standard linear schedule with T=1000:
| Step t | αˉt | αˉt (signal weight) | 1−αˉt (noise weight) | What it looks like |
|---|---|---|---|---|
| 0 | 1.000 | 1.000 | 0.000 | The original image |
| 100 | 0.897 | 0.947 | 0.321 | Slightly grainy; fully recognisable |
| 250 | 0.524 | 0.724 | 0.690 | Heavy grain; subject still identifiable |
| 500 | 0.079 | 0.281 | 0.960 | Mostly noise; vague blobs of colour |
| 750 | 0.003 | 0.058 | 0.998 | Effectively noise |
| 1000 | 0.00004 | 0.006 | 1.000 | Pure static |
Look at the middle of that table and a problem is visible. By step 500 the signal weight is already down to 0.28, and by 750 the image is gone. Half the schedule is spent on inputs that are indistinguishable from noise, teaching the network almost nothing. That observation is exactly what motivates the alternative schedule below.
What the network actually predicts
The obvious target is the clean image: give the network xt and have it output x0. It works, but it trains poorly. At high noise levels the clean image is nearly unrecoverable, so the errors are enormous and the gradients are dominated by the hardest examples.
The standard choice instead is to predict the noise ϵ that was added:
Three reasons this is better. The target always has the same statistics — unit-variance Gaussian noise — regardless of t, so the loss is on a consistent scale across the whole schedule. It is equivalent to predicting x0 up to a known rescaling, since x0=(xt−1−αˉtϵ)/αˉt, so nothing is lost. And it is a simple mean-squared-error objective with a known target: no adversarial game, no bound to be loose, just regression. This is why diffusion training is dramatically more reliable than GAN training.
An equivalent framing worth knowing: predicting the noise is, up to scaling, predicting ∇xlogp(xt) — the direction in which the data becomes more probable. Generation is then a walk uphill on the log-density, which is why this family is sometimes described as score-based.
Noise schedules
The choice of βt decides how the destruction is paced, and it matters more than its unglamorous name suggests.
| Schedule | Behaviour | Problem or benefit |
|---|---|---|
| Linear | β rises evenly from 10−4 to 0.02 | Destroys the image too early; the last third of steps are wasted. Fine at 256×256, poor at low resolution |
| Cosine | αˉt=cos2(1+st/T+s⋅2π) | Keeps meaningful signal much longer, then destroys it quickly at the end. Notably better sample quality |
| Sigmoid / shifted | Tunable midpoint | Higher resolutions need more total noise to erase structure; the schedule is shifted to compensate |
The cosine schedule's advantage is directly readable from the table above. Under the linear schedule αˉ has fallen to 0.08 by the midpoint; under cosine it is still around 0.5 there. The network therefore spends its training budget on noise levels where there is something left to learn.
The denoiser
The network sees a noisy image and a timestep, and outputs a noise prediction the same shape as the image. Two ingredients make that work.
Telling the network the time
The timestep is a single integer, and feeding a raw integer into a convolutional network is useless. Instead it is embedded with sinusoidal features — the same construction used for positions in a transformer — passed through a small multilayer perceptron, and then injected into every block, usually by predicting a per-channel scale and shift applied after normalisation.
1import math, torch23def timestep_embedding(t, dim):4 half = dim // 25 freqs = torch.exp(-math.log(10000) * torch.arange(half, device=t.device) / half)6 args = t[:, None].float() * freqs[None]7 return torch.cat([torch.cos(args), torch.sin(args)], dim=-1)This is not optional decoration. Without it the network cannot tell a nearly-clean image from a nearly-pure-noise one, and the correct behaviour in those two cases is completely different.
The backbone
Classically a U-Net: a contracting path that halves resolution while increasing channels, a bottleneck, and an expanding path back up, with skip connections joining matching resolutions. The skips are essential — the fine detail the model must eventually restore lives in the high-resolution early layers, and without a direct path it would have to survive the whole bottleneck. Self-attention blocks at the lower resolutions let distant regions coordinate, which is what keeps global composition coherent.
Training
1for x0 in loader:2 b = x0.size(0)3 t = torch.randint(0, T, (b,), device=dev) # a random step per example4 noise = torch.randn_like(x0)56 a_bar = alpha_bar[t].view(-1, 1, 1, 1)7 x_t = a_bar.sqrt() * x0 + (1 - a_bar).sqrt() * noise # closed-form jump89 loss = F.mse_loss(model(x_t, t), noise)10 opt.zero_grad(); loss.backward(); opt.step()Notice how ordinary this is. One network, one MSE loss, one optimiser. The random t per example is what makes it cheap: over a large dataset every noise level gets covered without any example ever being pushed through the full chain. There is no second network, no equilibrium, no bound. The loss curve descends and means something.
Sampling
Generation runs the chain backwards. Start from pure noise and repeatedly subtract a scaled version of the predicted noise, adding a little fresh noise back at each step except the last:
1x = torch.randn(n, 3, H, W, device=dev)2for t in reversed(range(T)):3 eps = model(x, torch.full((n,), t, device=dev))4 mean = (x - beta[t] / (1 - alpha_bar[t]).sqrt() * eps) / alpha[t].sqrt()5 x = mean + (sigma[t] * torch.randn_like(x) if t > 0 else 0)Faithful and slow: 1,000 network evaluations per image. Several routes cut that down.
| Sampler | Typical steps | Idea | Trade-off |
|---|---|---|---|
| DDPM (ancestral) | 1000 | Follow the derived reverse chain exactly | Highest fidelity, slowest |
| DDIM | 20–100 | A deterministic non-Markovian path that skips steps | Same noise gives the same image, enabling reproducibility and latent interpolation |
| DPM-Solver and relatives | 10–25 | Treat the reverse process as an ODE and use a higher-order solver | Excellent quality per step; more complex code |
| Distilled / consistency models | 1–4 | Train a student to reproduce the teacher's endpoint directly | Real-time generation; requires a separate distillation run and loses some quality |
DDIM's determinism is more useful than it first appears. Because a fixed starting noise always yields the same image, the initial noise becomes a reusable seed — you can interpolate between two seeds and get a smooth transition between two outputs, or edit a prompt while holding the seed fixed and see only the change you asked for.
Steering the output
Everything so far generates something plausible. To generate what you asked for, the condition — a text prompt, a class label, a depth map — has to enter the denoiser. Text is encoded into a sequence of vectors and injected through cross-attention layers, so at every denoising step every spatial location can look at every word.
Conditioning alone turns out to be too weak. The model follows the prompt loosely and drifts towards generic outputs. The fix that changed image generation in practice is classifier-free guidance.
During training, drop the condition around 10% of the time and replace it with a null token, so one network learns both the conditional and unconditional predictions. At sampling time, evaluate both and extrapolate away from the unconditional one:
The bracketed difference is the direction that the prompt adds. Multiplying it by w>1 pushes further along that direction than the model on its own would go — deliberately overshooting towards prompt adherence.
| Guidance w | Effect |
|---|---|
| 1.0 | Plain conditioning. Loose prompt adherence, high diversity |
| 3–5 | Balanced. Follows the prompt, keeps variety |
| 7–9 | The long-standing default for Stable Diffusion 1.x-era models. Strong adherence, reduced diversity |
| 15+ | Oversaturated colours, harsh contrast, visible artefacts |
The cost is doubled compute per step — two forward passes instead of one — and reduced diversity, since pushing every sample towards the prompt's centre necessarily narrows the spread. This is the dial that most affects perceived quality, and it is worth understanding rather than leaving at its default.
The right range depends on the model. Newer models often want much less: the published examples for Stable Diffusion 3.5 Large use 3.5 to 4.5, and FLUX.1 [dev] uses 3.5. Some fast models have guidance "distilled" into the network during training, so the number you pass is not the same knob and the doubled cost disappears. Start from the model card's value, not from a number that suited an older model.
Guidance does not make the model better at the prompt. It amplifies the difference the prompt already makes, and past a point that amplification becomes distortion.
Making it affordable
Two changes moved diffusion from research clusters to consumer hardware.
Latent diffusion. Running the process on 512×512×3 pixels means every one of dozens of steps touches 786,432 numbers. Instead, train an autoencoder that compresses images into a 64×64×4 latent grid, run the entire diffusion process there, and decode once at the end. That is 16,384 numbers per step — roughly 48 times fewer. The autoencoder handles the fine texture it is good at; diffusion handles the structure it is good at.
Diffusion transformers. Replacing the U-Net with a transformer over latent patches gives up convolution's built-in locality and gains predictable scaling: performance improves smoothly with parameters and compute in a way convolutional stacks do not match. Current large image and video systems are largely built this way, and the timestep and condition enter through modulated normalisation rather than through cross-attention alone.
Flow matching. Many of the newest models, including Stable Diffusion 3 and FLUX.1, keep the same recipe but change the training target. The noisy input is a straight-line blend, xt=(1−t)x0+tϵ with t running from 0 to 1, and the network predicts the velocity ϵ−x0: the direction from image to noise along that line. Sampling walks the line backwards by solving an ODE, exactly like the DPM-Solver row in the sampler table. Because the paths are straight rather than curved, they can be followed accurately in fewer steps. Everything else in this lesson, including the denoiser, latents, conditioning and guidance, carries over unchanged.
Diffusion against the alternatives
| Diffusion | GAN | VAE | |
|---|---|---|---|
| Sample quality | Best available | Sharp | Blurry |
| Mode coverage | Excellent | Unreliable | Good |
| Training stability | Plain MSE; very stable | Can fail outright | Very stable |
| Passes per sample | 10–1000 | 1 | 1 |
| Loss curve tells you quality | Roughly | No | Roughly |
| Controllability | Excellent — conditioning enters at every step | Moderate | Moderate |
| Editing an existing image | Natural — partially noise it, then denoise | Requires an inversion procedure | Encode and decode |
The editing row is underrated. Because generation passes through every noise level, you can enter the process partway: add noise to a real image up to step 400, then denoise from there with a new prompt. The result keeps the composition of the original and takes its content from the prompt, and the entry point is a continuous dial between "barely changed" and "completely new". Neither of the alternatives offers anything so direct.
What this means when you build something
Three knobs account for most of the difference between output that looks amateur and output that looks professional, and none of them requires retraining.
Steps. Below roughly 20 with a modern solver, quality degrades visibly. Above 50, improvements are usually imperceptible while cost rises linearly. Distilled models are the exception: they are built for 1 to 8 steps, and giving them more does not help. Test on your own prompts rather than trusting a default — the right number depends on the sampler, and the difference between 20 and 100 steps is a fivefold cost change.
Guidance. Start at the value on the model card and move it. If output looks washed out and ignores the prompt, raise it. If colours are lurid and contrast is crushed, lower it. Complex prompts generally want higher guidance; photographic realism generally wants lower.
The seed. With a deterministic sampler, fix the seed while you iterate on the prompt. Otherwise you cannot tell whether a change came from your edit or from a different random draw — which is the single most common way people waste hours convincing themselves a prompt trick works.
And when a project needs precise composition, stop trying to describe it. Structural conditioning — supplying an edge map, a depth map or a pose alongside the prompt — fixes the geometry and leaves only appearance to the model. That converts an unpredictable slot machine into a repeatable production tool, which is the difference between a demo and a pipeline.