Course Content
Image and Video Generation
4 sections · 7 lessons
Latent Diffusion Pipelines
Suppose you train a diffusion model the obvious way: directly on pixels. A 512×512 RGB image is 786,432 numbers, and the denoising network has to read and write all of them once per step, fifty steps per image. Worse, an attention layer over a 512×512 grid means a 262,144 × 262,144 attention matrix — 6.9×1010 entries, about 137 GB in half precision, per head. You are out of memory before the first forward pass finishes.
Pixel-space diffusion was done anyway — OpenAI's guided-diffusion used a cascade of a 64×64 model plus two super-resolution stages — and it cost hundreds of GPU-days. Latent diffusion cut that by roughly two orders of magnitude with one structural idea: most of those 786,432 numbers carry no perceptual information. Neighbouring pixels in a patch of sky are nearly identical; brick texture is statistically describable but not individually meaningful. Compress the image into something much smaller that preserves everything a human would notice, run diffusion there, and you get the same pictures for a fraction of the compute.
The compression step: what the VAE actually does
Stable Diffusion's encoder takes a 512×512×3 image and emits a 64×64×4 tensor: 16,384 numbers instead of 786,432, a 48× reduction. The grid shrinks 8× in each direction (hence "f8" autoencoder) and the 3 colour channels become 4 abstract ones with no direct visual meaning. A decoder runs it backwards. The pair is trained together on reconstruction long before any diffusion happens, then frozen — during generation the diffusion model never sees a pixel, and the decoder is called exactly once at the end.
Why "variational", and what the KL term buys you
A plain autoencoder would compress and reconstruct fine, but its latent space would be full of holes. Nothing stops the encoder scattering codes anywhere it likes, and the regions between valid codes decode to garbage. That is fatal here, because a diffusion model spends its entire life wandering the space, arriving at points no encoder ever produced.
A variational autoencoder fixes this. The encoder does not output a single latent; it outputs a mean and a variance, defining a small Gaussian blob. You sample from that blob to get the latent you decode. On top of the reconstruction loss you add a Kullback–Leibler penalty pulling each blob towards a standard normal distribution:
In plain English: reproduce the image accurately, but keep your encodings compact and centred near the origin rather than spread across the whole space. The two terms fight. Reconstruction wants precise, widely separated codes; the KL term wants everything packed into a smooth ball. The compromise is a latent space that is dense — nearby points decode to visually similar images, and points the encoder never emitted still decode to something plausible.
The KL term is not there to improve reconstruction. It is there so that the space between valid encodings is also valid, which is the only reason a diffusion model can navigate it at all.
Stable Diffusion sets the KL weight deliberately tiny (around 10−6): almost a plain autoencoder, sharp reconstruction, just enough regularisation to avoid a pathological space. The side effect is that latents are not unit-variance — they come out with standard deviation around 5.5. Diffusion training assumes roughly unit-scale data, so the pipeline multiplies encoder output by scale_factor = 0.18215 (approximately 1/5.489) and divides by it before decoding. Forget that constant and your images are saturated noise.
1import torch2from diffusers import AutoencoderKL34vae = AutoencoderKL.from_pretrained(5 "stabilityai/sd-vae-ft-mse", torch_dtype=torch.float166).to("cuda")78# image: a tensor of shape (1, 3, 512, 512) scaled to [-1, 1]9with torch.no_grad():10 posterior = vae.encode(image).latent_dist # mean + variance11 latent = posterior.sample() * vae.config.scaling_factor1213print(latent.shape) # torch.Size([1, 4, 64, 64])14print(latent.std().item()) # ~1.0 after scalingFailure mode people hit here: assuming the VAE is lossless. It is not. It reliably destroys small text, fine repeating patterns like chain-link fences, and detail in faces under about 100 pixels. No amount of prompt engineering fixes mushy text — the information was thrown away at 64×64 and the decoder is inventing plausible texture. This is also why upscaling then re-running at higher resolution helps: more latent pixels for the same content.
Forward diffusion: destroying a latent on a schedule
Training a diffusion model needs pairs of (corrupted input, known corruption). The forward process manufactures them. Starting from a clean latent z0, you repeatedly add a little Gaussian noise:
Read that as: at step t, shrink the current latent slightly by a factor of the square root of (1 minus beta), then add fresh Gaussian noise with variance beta. The shrink is what stops the values from growing without bound as noise accumulates.
Running that loop 1,000 times per training sample would be absurd. Fortunately Gaussians compose, and the whole chain collapses to a single closed-form jump. Define the cumulative product
In plain English: a noisy latent at any timestep is just a weighted blend of the clean latent and one draw of pure noise, with the weights fixed by the schedule. You can jump straight to t = 743 in one line of code. This is the single most important equation in the whole subject.
The numbers, for a linear schedule
Stable Diffusion 1.5 uses T = 1000 steps with β rising from 0.00085 to 0.012 on a "scaled linear" curve, linear in √β (the original DDPM paper used 0.0001 to 0.02; the arithmetic below uses the DDPM values because they are the canonical reference). Here is what the blend weights actually look like:
| t | βt | āt | signal weight √āt | noise weight √(1−āt) | SNR = ā/(1−ā) |
|---|---|---|---|---|---|
| 1 | 0.0001 | 0.9999 | 0.9999 | 0.0100 | 9999 |
| 100 | 0.0021 | 0.8970 | 0.9471 | 0.3209 | 8.71 |
| 200 | 0.0041 | 0.6590 | 0.8118 | 0.5839 | 1.93 |
| 400 | 0.0080 | 0.1951 | 0.4418 | 0.8971 | 0.24 |
| 600 | 0.0120 | 0.0259 | 0.1609 | 0.9870 | 0.027 |
| 800 | 0.0160 | 0.0015 | 0.0391 | 0.9992 | 0.0015 |
| 1000 | 0.0200 | 0.00004 | 0.0064 | 1.0000 | 0.00004 |
Work one element through. Take a single latent value z0 = 0.60 and a noise draw ε = −1.20.
- At t = 200: zt = 0.8118 × 0.60 + 0.5839 × (−1.20) = 0.4871 − 0.7007 = −0.2136. The original 0.60 has already flipped sign, but a good denoiser can still recover it — signal power is nearly twice noise power.
- At t = 600: zt = 0.1609 × 0.60 + 0.9870 × (−1.20) = 0.0965 − 1.1844 = −1.0879. The clean value contributes under 9% of the magnitude.
- At t = 1000: √ā = 0.0064, so z0 contributes 0.0038. Essentially pure noise.
That last row is why generation can start from scratch. If z1000 is statistically indistinguishable from torch.randn(1, 4, 64, 64), then sampling random noise and running the reverse chain lands you in the same distribution as real images.
Linear versus cosine schedules
Look again at the table. By t = 600 the SNR is 0.027 — the image is already gone. The remaining 400 timesteps, 40% of the training budget, teach the model to denoise noise into noise. That is wasted capacity.
The cosine schedule from Nichol and Dhariwal defines ā directly rather than defining β and multiplying up:
The small offset s stops βt being vanishingly small near t = 0; Nichol and Dhariwal found that such tiny amounts of early noise made ε hard to predict accurately. Compare the two directly:
| t | ā (linear) | ā (cosine) | What it means |
|---|---|---|---|
| 200 | 0.659 | 0.899 | Cosine destroys far less early — more steps spent on fine detail |
| 500 | 0.079 | 0.494 | At the halfway point, linear is done and cosine is still half signal |
| 800 | 0.0015 | 0.094 | Cosine still carries recoverable structure |
| 1000 | 0.00004 | ~0 | Both end at pure noise, as required |
A noise schedule is a budget allocation. It decides how many of your training steps are spent on coarse layout versus fine texture, and linear schedules spend far too many on regions where there is nothing left to learn.
Cosine helps most at low resolution, where linear destroys the image almost immediately. In 64×64 latent space the gap is smaller, which is why SD 1.5 shipped a rescaled linear schedule and got away with it.
What the network is actually predicting
This is where intuition usually goes wrong. The natural guess is that the network takes a noisy latent and outputs a cleaner latent. It does not. The standard formulation trains it to output the noise that was added.
In plain English: sample a clean latent, sample a noise vector, sample a timestep, mix them with the schedule weights, and ask the network to name the noise vector you used. Score it by mean squared error. That is the entire training objective. No adversarial loss, no perceptual loss, no discriminator.
Three things make noise-prediction the right target:
- It has constant scale. ε is standard normal at every timestep. Predicting z0 instead would make loss magnitudes incomparable across timesteps — near-trivial at t = 10, near-impossible at t = 900. ε gives a uniformly scaled target.
- It is algebraically equivalent to predicting the clean latent. Rearranging the forward equation: ẑ0 = (zt − √(1−āt) · εθ) / √āt. You can get one from the other for free. The difference is purely which one the loss is measured on, and therefore what the gradients emphasise.
- It is the score function in disguise. The gradient of the log-density of the noisy distribution is
In plain English: the predicted noise, scaled and negated, points in the direction that makes the current latent more probable. This is why the whole family is also called "score-based generative modelling", and it is the fact that makes classifier-free guidance work.
The v-prediction alternative
ε-prediction has one weak spot: at t = 1000, √ā is 0.0064, so recovering ẑ0 amplifies any prediction error by ~156×. SD 2.x's 768-pixel models and Stable Video Diffusion therefore use v-prediction, a rotated combination of both targets:
This target stays well-conditioned at both ends of the schedule. It is required if you want a true zero-terminal-SNR schedule — with āT exactly zero, ε-prediction becomes degenerate, because the input contains no information about the answer at all.
| Parameterisation | Target | Strong where | Used by |
|---|---|---|---|
| ε-prediction | the noise | mid-to-low timesteps | SD 1.x, SDXL base |
| x0-prediction | the clean latent | very high timesteps | rare alone; used inside samplers |
| v-prediction | rotated mix of both | uniformly, both ends | SD 2.x (768 px), Stable Video Diffusion |
| flow matching (velocity) | ε − x0 along a straight path | uniformly, both ends | SD 3.x, FLUX, Wan and many current video models |
The newest families have moved one step further. Stable Diffusion 3.5, FLUX.1 and FLUX.2, and open video models such as Wan 2.2 use flow matching (rectified flow): the noisy latent is a straight-line blend, zt = (1 − t)z0 + t·ε with t running from 0 to 1, and the network predicts the velocity ε − z0 along that line. It is the same idea as v-prediction — a target that stays well-behaved at both ends — with a simpler path. These models also replace the U-Net with a transformer. Everything else in this lesson still applies to them: a VAE latent, a noise level, a text-conditioned denoiser, guidance and a solver. SD 1.5 remains the clearest model to learn the mechanics on, which is why the examples use it.
One reverse step, worked in numbers
Generation starts from zT ~ N(0, I) and walks backwards. Take the deterministic DDIM update. We are at t = 500 moving to t' = 400; one latent element is zt = 0.83, and the network outputs εθ = 0.71.
Step 1 — estimate the clean latent. At t = 500, √ā = 0.2803 and √(1−ā) = 0.9599.
Step 2 — re-noise to the next timestep down. At t' = 400, √ā = 0.4418 and √(1−ā) = 0.8971. Reuse the same predicted noise:
Notice the shape of it. Each step makes a full guess at the final answer, then throws most of that guess away by re-adding noise appropriate to a slightly less noisy timestep. The guess at t = 500 is crude; by t = 100 it is nearly the finished image. The model never commits — it re-estimates from scratch every step, which is what lets it correct earlier mistakes.
Every denoising step predicts the entire final image, then discards most of that prediction. Sampling is 50 rough guesses that converge, not 50 incremental repairs.
Text conditioning: how the prompt gets in
Everything so far generates some image. To generate the one someone asked for, the prompt has to reach the network.
Prompt to vectors
CLIP's tokeniser pads or truncates the prompt to exactly 77 tokens, and the CLIP text encoder turns those into 77 vectors of dimension 768 (SD 1.5) or 1280+768 concatenated (SDXL, which runs two text encoders). The output is a 77×768 matrix — a sequence of contextualised word meanings, not one summary vector.
That limit is hard and it truncates silently. A 200-word prompt loses everything past roughly token 77 without warning — the single most common reason a carefully written prompt appears to be ignored.
Cross-attention
Inside the denoising network, at multiple resolutions, the image features query the text:
In plain English: every spatial position asks "which words are relevant to me?", gets a weighted answer, and mixes it in. Queries come from the image, keys and values from the prompt. At a 32×32 feature map that is 1,024 positions attending over 77 tokens — a 1,024×77 matrix per head, entirely affordable, unlike the pixel-space case above.
These maps are directly interpretable: visualise the attention row for the token "hat" and you get a heat map lighting up exactly where the model is placing a hat. Prompt-to-prompt editing works by editing those maps.
Classifier-free guidance, and why it works
Conditioning alone produces images only loosely related to the prompt. Ask for "a red bicycle against a blue wall" with no guidance and you get a pinkish bike, a wall drifting towards grey, and a general air of hedging. That is not a bug — the model is sampling honestly from a broad conditional distribution, and the mode of that distribution is mushy.
The fix — the most important inference-time trick in the field — is classifier-free guidance. During training the text conditioning is randomly dropped, replaced with an empty embedding, about 10% of the time. One network therefore learns both a conditional and an unconditional denoiser. At inference you run it twice and extrapolate:
In plain English: compute what the model would denoise towards with no prompt, compute what it would denoise towards with the prompt, and then move past the prompted answer in the direction that leads away from the unprompted one.
Why extrapolating is legitimate
Start from Bayes' rule on the noisy distribution. The conditional density factors as p(z|c) ∝ p(z) p(c|z). Take gradients of the log:
The second term is the score of an implicit classifier: "does this image look more like the prompt if I nudge it this way?" Sharpen that classifier by raising it to a power w, giving p(z)p(c∣z)w, whose score is
and since score is scaled negative predicted noise, that expression is the CFG formula. Guidance is not a hack: it is exact sampling from a distribution where the prompt-matching term has been exponentially sharpened. You are asking for images the implicit classifier is unusually confident about.
What the arithmetic does to your image
Take one latent element. The unconditional prediction is ε∅ = 0.42, the conditional is εc = 0.55.
| w | ε̂ = 0.42 + w(0.55 − 0.42) | Result | Typical output |
|---|---|---|---|
| 0 | 0.42 | ignores prompt entirely | a random plausible image |
| 1 | 0.55 | plain conditional sampling | on-topic but washed out, vague |
| 3 | 0.81 | moderate sharpening | natural, varied, good for photoreal |
| 7.5 | 1.395 | strong sharpening | the SD default; crisp, high contrast |
| 15 | 2.37 | far outside both predictions | blown highlights, neon colours, fried edges |
| 30 | 4.32 | wildly extrapolated | posterised sludge; often diverges |
At w = 7.5 the guided estimate is 1.395 — more than double either prediction the network actually made. That is the point and also the danger: you treat the difference between two model outputs as a direction and walk a long way along it, outside the region where the linear approximation holds. High CFG oversaturates because latents get pushed past the range the VAE decoder was trained on, and the decoder maps out-of-range latents to blown-out colour.
Cost: CFG doubles inference compute — every step runs the network twice. Both branches are batched into one forward pass of batch size 2, so it costs roughly 2× the FLOPs and 2× the activation memory.
Negative prompts are the same mechanism. Pass the embedding of "blurry, watermark, extra fingers" as the unconditional branch instead of the empty one, and the formula pushes actively away from that text rather than merely away from nothing. They cost nothing extra: that forward pass was already happening.
Assembling the pipeline
All five components now have jobs: tokeniser and text encoder produce a 77×768 matrix; the U-Net predicts noise from a latent, a timestep and that matrix; the scheduler steps between timesteps; the VAE decoder converts the final latent to pixels.
1import torch2from diffusers import StableDiffusionPipeline, DPMSolverMultistepScheduler34pipe = StableDiffusionPipeline.from_pretrained(5 "stable-diffusion-v1-5/stable-diffusion-v1-5", torch_dtype=torch.float166).to("cuda")7pipe.scheduler = DPMSolverMultistepScheduler.from_config(pipe.scheduler.config)89image = pipe(10 prompt="a red bicycle against a blue wall, golden hour",11 negative_prompt="blurry, watermark, low quality",12 num_inference_steps=25, guidance_scale=7.5,13 generator=torch.Generator("cuda").manual_seed(42),14).images[0]The manual version below is worth reading once, because everything discussed above appears in it literally:
1import torch23@torch.no_grad()4def embed(text):5 tok = pipe.tokenizer([text], padding="max_length", max_length=77,6 truncation=True, return_tensors="pt")7 return pipe.text_encoder(tok.input_ids.to("cuda"))[0]89embeddings = torch.cat([embed(""), embed("a red bicycle, blue wall")])1011latents = torch.randn((1, 4, 64, 64), device="cuda", dtype=torch.float16)12pipe.scheduler.set_timesteps(25)13latents = latents * pipe.scheduler.init_noise_sigma1415for t in pipe.scheduler.timesteps:16 inp = pipe.scheduler.scale_model_input(torch.cat([latents] * 2), t)17 with torch.no_grad():18 pred = pipe.unet(inp, t, encoder_hidden_states=embeddings).sample19 eps_uncond, eps_cond = pred.chunk(2) # CFG20 pred = eps_uncond + 7.5 * (eps_cond - eps_uncond)21 latents = pipe.scheduler.step(pred, t, latents).prev_sample2223with torch.no_grad(): # decode once, at the end24 image = pipe.vae.decode(latents / pipe.vae.config.scaling_factor).sampleSchedulers: how you walk the chain matters
Training uses 1,000 timesteps; sampling does not have to. A scheduler picks which subset to visit and how to integrate between them. The reverse process is a differential equation and schedulers are numerical solvers for it — better solvers need fewer evaluations for the same accuracy.
| Scheduler | Steps for good output | Deterministic? | Character | Use when |
|---|---|---|---|---|
| DDPM | ~1000 | no (adds noise each step) | the original; high diversity | reference / research only |
| DDIM | 50–100 | yes (η = 0) | reproducible; enables inversion | image editing, latent interpolation |
| Euler / Euler a | 20–30 | Euler yes, Euler-a no | fast, forgiving, ancestral variant is creative | general purpose, quick iteration |
| DPM-Solver++ (2M) | 15–25 | yes | best quality-per-step available | production default |
| LCM / Turbo | 1–8 | no (re-noises between steps) | distilled; needs a matching model | real-time and interactive tools |
Two traps. Swapping schedulers changes the image for a fixed seed, because different solvers follow different trajectories — a seed is not a portable identifier. And LCM and Turbo schedulers require distilled weights; pointing an LCM scheduler at stock SD 1.5 at 4 steps produces noise, and the failure looks like a scheduler bug rather than a model mismatch.
What this means when you build something
The architecture dictates where your problems will come from, and knowing the mechanism tells you which knob to reach for.
| Symptom | Actual cause | Fix |
|---|---|---|
| Text and small faces come out mushy | VAE compression discarded that detail at 64×64 | generate larger, or upscale then run img2img at low strength |
| Colours blown out, edges look fried | guidance scale pushed latents outside the decoder's trained range | drop guidance to 5–7, or enable guidance rescaling |
| End of a long prompt ignored | silent truncation at 77 CLIP tokens | shorten, front-load the important terms, or use a weighting library that chunks |
| Images look like grey noise | forgot the 0.18215 scale factor on encode or decode | multiply on encode, divide on decode |
| 4 steps produces garbage | LCM scheduler on non-distilled weights | load an LCM LoRA or a Turbo checkpoint, or raise steps to 20+ |
| Out of memory at 1024×1024 | attention memory grows quadratically with latent area | enable attention slicing and VAE tiling (pipe.vae.enable_tiling()); use enable_model_cpu_offload() |
On memory: float16 roughly halves weight memory and is near-free in quality for inference; enable_attention_slicing() computes attention in chunks, trading ~10% speed for a large saving; enable_model_cpu_offload() keeps only the currently-executing component on the GPU and gets SDXL into about 8 GB. That last one works precisely because the pipeline is five separable components running in sequence, not one monolith.
On cost: compute per step scales with latent area, so 512 to 1024 is 4× the latent pixels and more than 4× the attention cost. Halving step count by switching to DPM-Solver++ is usually a bigger and safer win than any other single change available to you.