Introduction to Generative AI

Diffusion Models


Take a photograph and add a small amount of random noise — enough to make it look slightly grainy. Now train a network to remove that grain. This is a solved problem, and has been for decades. It is ordinary regression: input a grainy image, output a clean one, minimise squared error. It trains in an afternoon and it works.

Now ask a bolder question. If a network can remove a little noise, what happens if you run it repeatedly? Start from an image that is 5% noise, denoise, and you have a clean image. Start from one that is 50% noise, denoise, denoise again, and again — does something emerge?

Push the idea to its limit. Start from pure noise, containing no image whatsoever, and apply the denoiser hundreds of times. Would a photograph appear out of static?

Remarkably, yes — but only if you fix one thing first. A denoiser trained on slightly grainy photographs has never seen pure static. Hand it a screen of random values and it produces nonsense, because that input is nothing like its training data. The repair is straightforward and it is the whole design: train one network across every noise level, from imperceptible grain to total static, and tell it at each call how noisy the input is supposed to be. Now it knows which job it is doing. At high noise it hallucinates coarse structure; at low noise it sharpens fine detail.

That is a diffusion model. Everything else is detail about how to add the noise, how to tell the network where it is, and how to take fewer steps.

The impossible leap from noise to a photograph becomes possible when it is broken into a thousand steps, each of which is a mild regression problem.

The forward schedule, in noise fraction0.000.150.420.710.931.00012345cleanimagepure noiseTraining picks one timestep at random and asks the network for the noise that was added there.
The hard problem of generation is cut into hundreds of easy denoising problems, each one plain regression.

The forward process: destroying an image on purpose

Define a fixed recipe for corruption. At each step tt, shrink the current image slightly and add a little Gaussian noise:

xt=1−βt xt−1+βt ϵ,ϵ∼N(0,I)x_t = \sqrt{1 - \beta_t}\, x_{t-1} + \sqrt{\beta_t}\, \epsilon, \qquad \epsilon \sim \mathcal{N}(0, I)

The βt\beta_t values are small — typically from 10−410^{-4} up to 0.020.02 over T=1000T = 1000 steps — and are chosen in advance, not learned. Run this to completion and any image becomes indistinguishable from pure noise.

Simulating a thousand steps for every training example would be crippling, so the key practical fact is that this process has a closed form. Writing αt=1−βt\alpha_t = 1 - \beta_t and αˉt=∏s=1tαs\bar{\alpha}_t = \prod_{s=1}^{t} \alpha_s:

xt=αˉt x0+1−αˉt ϵx_t = \sqrt{\bar{\alpha}_t}\, x_0 + \sqrt{1 - \bar{\alpha}_t}\, \epsilon

One line, one noise draw, and you jump directly to any step. αˉt\bar{\alpha}_t is the fraction of the original signal that survives; 1−αˉt1 - \bar{\alpha}_t is the noise fraction. They sum to one, which is why the image's overall scale stays stable rather than exploding.

Concrete values for the standard linear schedule with T=1000T = 1000:

Step ttαˉt\bar{\alpha}_tαˉt\sqrt{\bar{\alpha}_t} (signal weight)1−αˉt\sqrt{1-\bar{\alpha}_t} (noise weight)What it looks like
01.0001.0000.000The original image
1000.8970.9470.321Slightly grainy; fully recognisable
2500.5240.7240.690Heavy grain; subject still identifiable
5000.0790.2810.960Mostly noise; vague blobs of colour
7500.0030.0580.998Effectively noise
10000.000040.0061.000Pure static

Look at the middle of that table and a problem is visible. By step 500 the signal weight is already down to 0.28, and by 750 the image is gone. Half the schedule is spent on inputs that are indistinguishable from noise, teaching the network almost nothing. That observation is exactly what motivates the alternative schedule below.

What the network actually predicts

The obvious target is the clean image: give the network xtx_t and have it output x0x_0. It works, but it trains poorly. At high noise levels the clean image is nearly unrecoverable, so the errors are enormous and the gradients are dominated by the hardest examples.

The standard choice instead is to predict the noise ϵ\epsilon that was added:

L=Ex0, ϵ, t[∥ϵ−ϵθ(xt,t)∥2]\mathcal{L} = \mathbb{E}_{x_0,\, \epsilon,\, t}\left[\big\| \epsilon - \epsilon_\theta(x_t, t) \big\|^2\right]

Three reasons this is better. The target always has the same statistics — unit-variance Gaussian noise — regardless of tt, so the loss is on a consistent scale across the whole schedule. It is equivalent to predicting x0x_0 up to a known rescaling, since x0=(xt−1−αˉt ϵ)/αˉtx_0 = (x_t - \sqrt{1-\bar\alpha_t}\,\epsilon)/\sqrt{\bar\alpha_t}, so nothing is lost. And it is a simple mean-squared-error objective with a known target: no adversarial game, no bound to be loose, just regression. This is why diffusion training is dramatically more reliable than GAN training.

An equivalent framing worth knowing: predicting the noise is, up to scaling, predicting ∇xlog⁡p(xt)\nabla_{x} \log p(x_t) — the direction in which the data becomes more probable. Generation is then a walk uphill on the log-density, which is why this family is sometimes described as score-based.

Noise schedules

The choice of βt\beta_t decides how the destruction is paced, and it matters more than its unglamorous name suggests.

ScheduleBehaviourProblem or benefit
Linearβ\beta rises evenly from 10−410^{-4} to 0.020.02Destroys the image too early; the last third of steps are wasted. Fine at 256×256, poor at low resolution
Cosineαˉt=cos⁡2 ⁣(t/T+s1+s⋅π2)\bar{\alpha}_t = \cos^2\!\left(\frac{t/T + s}{1 + s} \cdot \frac{\pi}{2}\right)Keeps meaningful signal much longer, then destroys it quickly at the end. Notably better sample quality
Sigmoid / shiftedTunable midpointHigher resolutions need more total noise to erase structure; the schedule is shifted to compensate

The cosine schedule's advantage is directly readable from the table above. Under the linear schedule αˉ\bar\alpha has fallen to 0.08 by the midpoint; under cosine it is still around 0.5 there. The network therefore spends its training budget on noise levels where there is something left to learn.

The denoiser

The network sees a noisy image and a timestep, and outputs a noise prediction the same shape as the image. Two ingredients make that work.

Telling the network the time

The timestep is a single integer, and feeding a raw integer into a convolutional network is useless. Instead it is embedded with sinusoidal features — the same construction used for positions in a transformer — passed through a small multilayer perceptron, and then injected into every block, usually by predicting a per-channel scale and shift applied after normalisation.

Python
import math, torchdef timestep_embedding(t, dim):    half = dim // 2    freqs = torch.exp(-math.log(10000) * torch.arange(half, device=t.device) / half)    args  = t[:, None].float() * freqs[None]    return torch.cat([torch.cos(args), torch.sin(args)], dim=-1)

This is not optional decoration. Without it the network cannot tell a nearly-clean image from a nearly-pure-noise one, and the correct behaviour in those two cases is completely different.

The backbone

Classically a U-Net: a contracting path that halves resolution while increasing channels, a bottleneck, and an expanding path back up, with skip connections joining matching resolutions. The skips are essential — the fine detail the model must eventually restore lives in the high-resolution early layers, and without a direct path it would have to survive the whole bottleneck. Self-attention blocks at the lower resolutions let distant regions coordinate, which is what keeps global composition coherent.

Training

Python
for x0 in loader:    b = x0.size(0)    t = torch.randint(0, T, (b,), device=dev)          # a random step per example    noise = torch.randn_like(x0)    a_bar = alpha_bar[t].view(-1, 1, 1, 1)    x_t   = a_bar.sqrt() * x0 + (1 - a_bar).sqrt() * noise   # closed-form jump    loss = F.mse_loss(model(x_t, t), noise)    opt.zero_grad(); loss.backward(); opt.step()

Notice how ordinary this is. One network, one MSE loss, one optimiser. The random tt per example is what makes it cheap: over a large dataset every noise level gets covered without any example ever being pushed through the full chain. There is no second network, no equilibrium, no bound. The loss curve descends and means something.

Sampling

Generation runs the chain backwards. Start from pure noise and repeatedly subtract a scaled version of the predicted noise, adding a little fresh noise back at each step except the last:

Python
x = torch.randn(n, 3, H, W, device=dev)for t in reversed(range(T)):    eps = model(x, torch.full((n,), t, device=dev))    mean = (x - beta[t] / (1 - alpha_bar[t]).sqrt() * eps) / alpha[t].sqrt()    x = mean + (sigma[t] * torch.randn_like(x) if t > 0 else 0)

Faithful and slow: 1,000 network evaluations per image. Several routes cut that down.

SamplerTypical stepsIdeaTrade-off
DDPM (ancestral)1000Follow the derived reverse chain exactlyHighest fidelity, slowest
DDIM20–100A deterministic non-Markovian path that skips stepsSame noise gives the same image, enabling reproducibility and latent interpolation
DPM-Solver and relatives10–25Treat the reverse process as an ODE and use a higher-order solverExcellent quality per step; more complex code
Distilled / consistency models1–4Train a student to reproduce the teacher's endpoint directlyReal-time generation; requires a separate distillation run and loses some quality

DDIM's determinism is more useful than it first appears. Because a fixed starting noise always yields the same image, the initial noise becomes a reusable seed — you can interpolate between two seeds and get a smooth transition between two outputs, or edit a prompt while holding the seed fixed and see only the change you asked for.

Steering the output

Everything so far generates something plausible. To generate what you asked for, the condition — a text prompt, a class label, a depth map — has to enter the denoiser. Text is encoded into a sequence of vectors and injected through cross-attention layers, so at every denoising step every spatial location can look at every word.

Conditioning alone turns out to be too weak. The model follows the prompt loosely and drifts towards generic outputs. The fix that changed image generation in practice is classifier-free guidance.

During training, drop the condition around 10% of the time and replace it with a null token, so one network learns both the conditional and unconditional predictions. At sampling time, evaluate both and extrapolate away from the unconditional one:

ϵ~=ϵθ(xt,∅)+w⋅(ϵθ(xt,c)−ϵθ(xt,∅))\tilde{\epsilon} = \epsilon_\theta(x_t, \varnothing) + w \cdot \big(\epsilon_\theta(x_t, c) - \epsilon_\theta(x_t, \varnothing)\big)

The bracketed difference is the direction that the prompt adds. Multiplying it by w>1w > 1 pushes further along that direction than the model on its own would go — deliberately overshooting towards prompt adherence.

Guidance wwEffect
1.0Plain conditioning. Loose prompt adherence, high diversity
3–5Balanced. Follows the prompt, keeps variety
7–9The long-standing default for Stable Diffusion 1.x-era models. Strong adherence, reduced diversity
15+Oversaturated colours, harsh contrast, visible artefacts

The cost is doubled compute per step — two forward passes instead of one — and reduced diversity, since pushing every sample towards the prompt's centre necessarily narrows the spread. This is the dial that most affects perceived quality, and it is worth understanding rather than leaving at its default.

The right range depends on the model. Newer models often want much less: the published examples for Stable Diffusion 3.5 Large use 3.5 to 4.5, and FLUX.1 [dev] uses 3.5. Some fast models have guidance "distilled" into the network during training, so the number you pass is not the same knob and the doubled cost disappears. Start from the model card's value, not from a number that suited an older model.

Guidance does not make the model better at the prompt. It amplifies the difference the prompt already makes, and past a point that amplification becomes distortion.

Making it affordable

Two changes moved diffusion from research clusters to consumer hardware.

Latent diffusion. Running the process on 512×512×3 pixels means every one of dozens of steps touches 786,432 numbers. Instead, train an autoencoder that compresses images into a 64×64×4 latent grid, run the entire diffusion process there, and decode once at the end. That is 16,384 numbers per step — roughly 48 times fewer. The autoencoder handles the fine texture it is good at; diffusion handles the structure it is good at.

Diffusion transformers. Replacing the U-Net with a transformer over latent patches gives up convolution's built-in locality and gains predictable scaling: performance improves smoothly with parameters and compute in a way convolutional stacks do not match. Current large image and video systems are largely built this way, and the timestep and condition enter through modulated normalisation rather than through cross-attention alone.

Flow matching. Many of the newest models, including Stable Diffusion 3 and FLUX.1, keep the same recipe but change the training target. The noisy input is a straight-line blend, xt=(1−t) x0+t ϵx_t = (1 - t)\, x_0 + t\, \epsilon with tt running from 0 to 1, and the network predicts the velocity ϵ−x0\epsilon - x_0: the direction from image to noise along that line. Sampling walks the line backwards by solving an ODE, exactly like the DPM-Solver row in the sampler table. Because the paths are straight rather than curved, they can be followed accurately in fewer steps. Everything else in this lesson, including the denoiser, latents, conditioning and guidance, carries over unchanged.

Diffusion against the alternatives

DiffusionGANVAE
Sample qualityBest availableSharpBlurry
Mode coverageExcellentUnreliableGood
Training stabilityPlain MSE; very stableCan fail outrightVery stable
Passes per sample10–100011
Loss curve tells you qualityRoughlyNoRoughly
ControllabilityExcellent — conditioning enters at every stepModerateModerate
Editing an existing imageNatural — partially noise it, then denoiseRequires an inversion procedureEncode and decode

The editing row is underrated. Because generation passes through every noise level, you can enter the process partway: add noise to a real image up to step 400, then denoise from there with a new prompt. The result keeps the composition of the original and takes its content from the prompt, and the entry point is a continuous dial between "barely changed" and "completely new". Neither of the alternatives offers anything so direct.

What this means when you build something

Three knobs account for most of the difference between output that looks amateur and output that looks professional, and none of them requires retraining.

Steps. Below roughly 20 with a modern solver, quality degrades visibly. Above 50, improvements are usually imperceptible while cost rises linearly. Distilled models are the exception: they are built for 1 to 8 steps, and giving them more does not help. Test on your own prompts rather than trusting a default — the right number depends on the sampler, and the difference between 20 and 100 steps is a fivefold cost change.

Guidance. Start at the value on the model card and move it. If output looks washed out and ignores the prompt, raise it. If colours are lurid and contrast is crushed, lower it. Complex prompts generally want higher guidance; photographic realism generally wants lower.

The seed. With a deterministic sampler, fix the seed while you iterate on the prompt. Otherwise you cannot tell whether a change came from your edit or from a different random draw — which is the single most common way people waste hours convincing themselves a prompt trick works.

And when a project needs precise composition, stop trying to describe it. Structural conditioning — supplying an edge map, a depth map or a pose alongside the prompt — fixes the geometry and leaves only appearance to the model. That converts an unpredictable slot machine into a repeatable production tool, which is the difference between a demo and a pipeline.