Image and Video Generation

Latent Diffusion Pipelines


Suppose you train a diffusion model the obvious way: directly on pixels. A 512×512 RGB image is 786,432 numbers, and the denoising network has to read and write all of them once per step, fifty steps per image. Worse, an attention layer over a 512×512 grid means a 262,144 × 262,144 attention matrix — 6.9×10106.9 \times 10^{10} entries, about 137 GB in half precision, per head. You are out of memory before the first forward pass finishes.

Pixel-space diffusion was done anyway — OpenAI's guided-diffusion used a cascade of a 64×64 model plus two super-resolution stages — and it cost hundreds of GPU-days. Latent diffusion cut that by roughly two orders of magnitude with one structural idea: most of those 786,432 numbers carry no perceptual information. Neighbouring pixels in a patch of sky are nearly identical; brick texture is statistically describable but not individually meaningful. Compress the image into something much smaller that preserves everything a human would notice, run diffusion there, and you get the same pictures for a fraction of the compute.

Why diffusion moved out of pixel spaceImage: 512 by512 by 3 =786,432 numbersVAE encodercompresses 48 to 1Latent: 64 by64 by 4 =16,384 numbersDenoise here,25 to 50 stepsVAE decoderpaints thepixels backAttention over a 512 by 512 grid needs a 262,144 square matrix; over a 64 by 64 latent it is 4,096 square.
The VAE does the perceptual work once, so the expensive iterative loop runs on a grid 48 times smaller.

The compression step: what the VAE actually does

Stable Diffusion's encoder takes a 512×512×3 image and emits a 64×64×4 tensor: 16,384 numbers instead of 786,432, a 48× reduction. The grid shrinks 8× in each direction (hence "f8" autoencoder) and the 3 colour channels become 4 abstract ones with no direct visual meaning. A decoder runs it backwards. The pair is trained together on reconstruction long before any diffusion happens, then frozen — during generation the diffusion model never sees a pixel, and the decoder is called exactly once at the end.

Why "variational", and what the KL term buys you

A plain autoencoder would compress and reconstruct fine, but its latent space would be full of holes. Nothing stops the encoder scattering codes anywhere it likes, and the regions between valid codes decode to garbage. That is fatal here, because a diffusion model spends its entire life wandering the space, arriving at points no encoder ever produced.

A variational autoencoder fixes this. The encoder does not output a single latent; it outputs a mean and a variance, defining a small Gaussian blob. You sample from that blob to get the latent you decode. On top of the reconstruction loss you add a Kullback–Leibler penalty pulling each blob towards a standard normal distribution:

LVAE=∥x−x^∥2⏟reconstruction+  λ⋅DKL ⁣(q(z∣x) ∥ N(0,I))⏟regularisation\mathcal{L}_{\text{VAE}} = \underbrace{\lVert x - \hat{x}\rVert^2}_{\text{reconstruction}} + \; \lambda \cdot \underbrace{D_{\mathrm{KL}}\!\left(q(z \mid x)\,\Vert\,\mathcal{N}(0, I)\right)}_{\text{regularisation}}

In plain English: reproduce the image accurately, but keep your encodings compact and centred near the origin rather than spread across the whole space. The two terms fight. Reconstruction wants precise, widely separated codes; the KL term wants everything packed into a smooth ball. The compromise is a latent space that is dense — nearby points decode to visually similar images, and points the encoder never emitted still decode to something plausible.

The KL term is not there to improve reconstruction. It is there so that the space between valid encodings is also valid, which is the only reason a diffusion model can navigate it at all.

Stable Diffusion sets the KL weight deliberately tiny (around 10−610^{-6}): almost a plain autoencoder, sharp reconstruction, just enough regularisation to avoid a pathological space. The side effect is that latents are not unit-variance — they come out with standard deviation around 5.5. Diffusion training assumes roughly unit-scale data, so the pipeline multiplies encoder output by scale_factor = 0.18215 (approximately 1/5.489) and divides by it before decoding. Forget that constant and your images are saturated noise.

Python
import torchfrom diffusers import AutoencoderKLvae = AutoencoderKL.from_pretrained(    "stabilityai/sd-vae-ft-mse", torch_dtype=torch.float16).to("cuda")# image: a tensor of shape (1, 3, 512, 512) scaled to [-1, 1]with torch.no_grad():    posterior = vae.encode(image).latent_dist   # mean + variance    latent = posterior.sample() * vae.config.scaling_factorprint(latent.shape)        # torch.Size([1, 4, 64, 64])print(latent.std().item()) # ~1.0 after scaling

Failure mode people hit here: assuming the VAE is lossless. It is not. It reliably destroys small text, fine repeating patterns like chain-link fences, and detail in faces under about 100 pixels. No amount of prompt engineering fixes mushy text — the information was thrown away at 64×64 and the decoder is inventing plausible texture. This is also why upscaling then re-running at higher resolution helps: more latent pixels for the same content.

Forward diffusion: destroying a latent on a schedule

Training a diffusion model needs pairs of (corrupted input, known corruption). The forward process manufactures them. Starting from a clean latent z0, you repeatedly add a little Gaussian noise:

q(zt∣zt−1)=N ⁣(zt; 1−βt zt−1, βtI)q(z_t \mid z_{t-1}) = \mathcal{N}\!\left(z_t;\ \sqrt{1-\beta_t}\, z_{t-1},\ \beta_t I\right)

Read that as: at step t, shrink the current latent slightly by a factor of the square root of (1 minus beta), then add fresh Gaussian noise with variance beta. The shrink is what stops the values from growing without bound as noise accumulates.

Running that loop 1,000 times per training sample would be absurd. Fortunately Gaussians compose, and the whole chain collapses to a single closed-form jump. Define the cumulative product

αˉt=∏s=1t(1−βs),thenzt=αˉt z0+1−αˉt ϵ,ϵ∼N(0,I)\bar{\alpha}_t = \prod_{s=1}^{t}(1 - \beta_s), \qquad\text{then}\qquad z_t = \sqrt{\bar{\alpha}_t}\, z_0 + \sqrt{1 - \bar{\alpha}_t}\,\epsilon, \quad \epsilon \sim \mathcal{N}(0, I)

In plain English: a noisy latent at any timestep is just a weighted blend of the clean latent and one draw of pure noise, with the weights fixed by the schedule. You can jump straight to t = 743 in one line of code. This is the single most important equation in the whole subject.

The numbers, for a linear schedule

Stable Diffusion 1.5 uses T = 1000 steps with β rising from 0.00085 to 0.012 on a "scaled linear" curve, linear in √β (the original DDPM paper used 0.0001 to 0.02; the arithmetic below uses the DDPM values because they are the canonical reference). Here is what the blend weights actually look like:

tβtātsignal weight √ātnoise weight √(1−āt)SNR = ā/(1−ā)
10.00010.99990.99990.01009999
1000.00210.89700.94710.32098.71
2000.00410.65900.81180.58391.93
4000.00800.19510.44180.89710.24
6000.01200.02590.16090.98700.027
8000.01600.00150.03910.99920.0015
10000.02000.000040.00641.00000.00004

Work one element through. Take a single latent value z0 = 0.60 and a noise draw ε = −1.20.

  • At t = 200: zt = 0.8118 × 0.60 + 0.5839 × (−1.20) = 0.4871 − 0.7007 = −0.2136. The original 0.60 has already flipped sign, but a good denoiser can still recover it — signal power is nearly twice noise power.
  • At t = 600: zt = 0.1609 × 0.60 + 0.9870 × (−1.20) = 0.0965 − 1.1844 = −1.0879. The clean value contributes under 9% of the magnitude.
  • At t = 1000: √ā = 0.0064, so z0 contributes 0.0038. Essentially pure noise.

That last row is why generation can start from scratch. If z1000 is statistically indistinguishable from torch.randn(1, 4, 64, 64), then sampling random noise and running the reverse chain lands you in the same distribution as real images.

Linear versus cosine schedules

Look again at the table. By t = 600 the SNR is 0.027 — the image is already gone. The remaining 400 timesteps, 40% of the training budget, teach the model to denoise noise into noise. That is wasted capacity.

The cosine schedule from Nichol and Dhariwal defines ā directly rather than defining β and multiplying up:

αˉt=f(t)f(0),f(t)=cos⁡2 ⁣(t/T+s1+s⋅π2),s=0.008\bar{\alpha}_t = \frac{f(t)}{f(0)}, \qquad f(t) = \cos^2\!\left(\frac{t/T + s}{1 + s}\cdot\frac{\pi}{2}\right), \quad s = 0.008

The small offset s stops βt being vanishingly small near t = 0; Nichol and Dhariwal found that such tiny amounts of early noise made ε hard to predict accurately. Compare the two directly:

tā (linear)ā (cosine)What it means
2000.6590.899Cosine destroys far less early — more steps spent on fine detail
5000.0790.494At the halfway point, linear is done and cosine is still half signal
8000.00150.094Cosine still carries recoverable structure
10000.00004~0Both end at pure noise, as required

A noise schedule is a budget allocation. It decides how many of your training steps are spent on coarse layout versus fine texture, and linear schedules spend far too many on regions where there is nothing left to learn.

Cosine helps most at low resolution, where linear destroys the image almost immediately. In 64×64 latent space the gap is smaller, which is why SD 1.5 shipped a rescaled linear schedule and got away with it.

What the network is actually predicting

This is where intuition usually goes wrong. The natural guess is that the network takes a noisy latent and outputs a cleaner latent. It does not. The standard formulation trains it to output the noise that was added.

L=Ez0, ϵ, t[∥ϵ−ϵθ(zt,t,c)∥2]\mathcal{L} = \mathbb{E}_{z_0,\,\epsilon,\,t}\left[\left\lVert \epsilon - \epsilon_\theta(z_t, t, c)\right\rVert^2\right]

In plain English: sample a clean latent, sample a noise vector, sample a timestep, mix them with the schedule weights, and ask the network to name the noise vector you used. Score it by mean squared error. That is the entire training objective. No adversarial loss, no perceptual loss, no discriminator.

Three things make noise-prediction the right target:

  1. It has constant scale. ε is standard normal at every timestep. Predicting z0 instead would make loss magnitudes incomparable across timesteps — near-trivial at t = 10, near-impossible at t = 900. ε gives a uniformly scaled target.
  2. It is algebraically equivalent to predicting the clean latent. Rearranging the forward equation: ẑ0 = (zt − √(1−āt) · εθ) / √āt. You can get one from the other for free. The difference is purely which one the loss is measured on, and therefore what the gradients emphasise.
  3. It is the score function in disguise. The gradient of the log-density of the noisy distribution is

∇ztlog⁡pt(zt)=−ϵθ(zt,t)1−αˉt\nabla_{z_t} \log p_t(z_t) = -\frac{\epsilon_\theta(z_t, t)}{\sqrt{1 - \bar{\alpha}_t}}

In plain English: the predicted noise, scaled and negated, points in the direction that makes the current latent more probable. This is why the whole family is also called "score-based generative modelling", and it is the fact that makes classifier-free guidance work.

The v-prediction alternative

ε-prediction has one weak spot: at t = 1000, √ā is 0.0064, so recovering ẑ0 amplifies any prediction error by ~156×. SD 2.x's 768-pixel models and Stable Video Diffusion therefore use v-prediction, a rotated combination of both targets:

vt=αˉt ϵ−1−αˉt z0v_t = \sqrt{\bar{\alpha}_t}\,\epsilon - \sqrt{1 - \bar{\alpha}_t}\,z_0

This target stays well-conditioned at both ends of the schedule. It is required if you want a true zero-terminal-SNR schedule — with āT exactly zero, ε-prediction becomes degenerate, because the input contains no information about the answer at all.

ParameterisationTargetStrong whereUsed by
ε-predictionthe noisemid-to-low timestepsSD 1.x, SDXL base
x0-predictionthe clean latentvery high timestepsrare alone; used inside samplers
v-predictionrotated mix of bothuniformly, both endsSD 2.x (768 px), Stable Video Diffusion
flow matching (velocity)ε − x0 along a straight pathuniformly, both endsSD 3.x, FLUX, Wan and many current video models

The newest families have moved one step further. Stable Diffusion 3.5, FLUX.1 and FLUX.2, and open video models such as Wan 2.2 use flow matching (rectified flow): the noisy latent is a straight-line blend, zt = (1 − t)z0 + t·ε with t running from 0 to 1, and the network predicts the velocity ε − z0 along that line. It is the same idea as v-prediction — a target that stays well-behaved at both ends — with a simpler path. These models also replace the U-Net with a transformer. Everything else in this lesson still applies to them: a VAE latent, a noise level, a text-conditioned denoiser, guidance and a solver. SD 1.5 remains the clearest model to learn the mechanics on, which is why the examples use it.

One reverse step, worked in numbers

Generation starts from zT ~ N(0, I) and walks backwards. Take the deterministic DDIM update. We are at t = 500 moving to t' = 400; one latent element is zt = 0.83, and the network outputs εθ = 0.71.

Step 1 — estimate the clean latent. At t = 500, √ā = 0.2803 and √(1−ā) = 0.9599.

z^0=0.83−0.9599×0.710.2803=0.83−0.68150.2803=0.14850.2803=0.530\hat{z}_0 = \frac{0.83 - 0.9599 \times 0.71}{0.2803} = \frac{0.83 - 0.6815}{0.2803} = \frac{0.1485}{0.2803} = 0.530

Step 2 — re-noise to the next timestep down. At t' = 400, √ā = 0.4418 and √(1−ā) = 0.8971. Reuse the same predicted noise:

zt′=0.4418×0.530+0.8971×0.71=0.2340+0.6369=0.871z_{t'} = 0.4418 \times 0.530 + 0.8971 \times 0.71 = 0.2340 + 0.6369 = 0.871

Notice the shape of it. Each step makes a full guess at the final answer, then throws most of that guess away by re-adding noise appropriate to a slightly less noisy timestep. The guess at t = 500 is crude; by t = 100 it is nearly the finished image. The model never commits — it re-estimates from scratch every step, which is what lets it correct earlier mistakes.

Every denoising step predicts the entire final image, then discards most of that prediction. Sampling is 50 rough guesses that converge, not 50 incremental repairs.

Text conditioning: how the prompt gets in

Everything so far generates some image. To generate the one someone asked for, the prompt has to reach the network.

Prompt to vectors

CLIP's tokeniser pads or truncates the prompt to exactly 77 tokens, and the CLIP text encoder turns those into 77 vectors of dimension 768 (SD 1.5) or 1280+768 concatenated (SDXL, which runs two text encoders). The output is a 77×768 matrix — a sequence of contextualised word meanings, not one summary vector.

That limit is hard and it truncates silently. A 200-word prompt loses everything past roughly token 77 without warning — the single most common reason a carefully written prompt appears to be ignored.

Cross-attention

Inside the denoising network, at multiple resolutions, the image features query the text:

Attention(Q,K,V)=softmax ⁣(QK⊤dk)V,Q=WQ himage,K=WK ctext,V=WV ctext\text{Attention}(Q, K, V) = \mathrm{softmax}\!\left(\frac{QK^{\top}}{\sqrt{d_k}}\right)V, \qquad Q = W_Q\,h_{\text{image}},\quad K = W_K\,c_{\text{text}},\quad V = W_V\,c_{\text{text}}

In plain English: every spatial position asks "which words are relevant to me?", gets a weighted answer, and mixes it in. Queries come from the image, keys and values from the prompt. At a 32×32 feature map that is 1,024 positions attending over 77 tokens — a 1,024×77 matrix per head, entirely affordable, unlike the pixel-space case above.

These maps are directly interpretable: visualise the attention row for the token "hat" and you get a heat map lighting up exactly where the model is placing a hat. Prompt-to-prompt editing works by editing those maps.

Classifier-free guidance, and why it works

Conditioning alone produces images only loosely related to the prompt. Ask for "a red bicycle against a blue wall" with no guidance and you get a pinkish bike, a wall drifting towards grey, and a general air of hedging. That is not a bug — the model is sampling honestly from a broad conditional distribution, and the mode of that distribution is mushy.

The fix — the most important inference-time trick in the field — is classifier-free guidance. During training the text conditioning is randomly dropped, replaced with an empty embedding, about 10% of the time. One network therefore learns both a conditional and an unconditional denoiser. At inference you run it twice and extrapolate:

ϵ^=ϵθ(zt,∅)+w[ϵθ(zt,c)−ϵθ(zt,∅)]\hat{\epsilon} = \epsilon_\theta(z_t, \varnothing) + w\left[\epsilon_\theta(z_t, c) - \epsilon_\theta(z_t, \varnothing)\right]

In plain English: compute what the model would denoise towards with no prompt, compute what it would denoise towards with the prompt, and then move past the prompted answer in the direction that leads away from the unprompted one.

Why extrapolating is legitimate

Start from Bayes' rule on the noisy distribution. The conditional density factors as p(z|c) ∝ p(z) p(c|z). Take gradients of the log:

∇zlog⁡p(z∣c)=∇zlog⁡p(z)+∇zlog⁡p(c∣z)\nabla_z \log p(z \mid c) = \nabla_z \log p(z) + \nabla_z \log p(c \mid z)

The second term is the score of an implicit classifier: "does this image look more like the prompt if I nudge it this way?" Sharpen that classifier by raising it to a power w, giving p(z) p(c∣z)wp(z)\,p(c \mid z)^{w}, whose score is

∇zlog⁡p(z)+w[∇zlog⁡p(z∣c)−∇zlog⁡p(z)]\nabla_z \log p(z) + w\left[\nabla_z \log p(z \mid c) - \nabla_z \log p(z)\right]

and since score is scaled negative predicted noise, that expression is the CFG formula. Guidance is not a hack: it is exact sampling from a distribution where the prompt-matching term has been exponentially sharpened. You are asking for images the implicit classifier is unusually confident about.

What the arithmetic does to your image

Take one latent element. The unconditional prediction is ε∅ = 0.42, the conditional is εc = 0.55.

wε̂ = 0.42 + w(0.55 − 0.42)ResultTypical output
00.42ignores prompt entirelya random plausible image
10.55plain conditional samplingon-topic but washed out, vague
30.81moderate sharpeningnatural, varied, good for photoreal
7.51.395strong sharpeningthe SD default; crisp, high contrast
152.37far outside both predictionsblown highlights, neon colours, fried edges
304.32wildly extrapolatedposterised sludge; often diverges

At w = 7.5 the guided estimate is 1.395 — more than double either prediction the network actually made. That is the point and also the danger: you treat the difference between two model outputs as a direction and walk a long way along it, outside the region where the linear approximation holds. High CFG oversaturates because latents get pushed past the range the VAE decoder was trained on, and the decoder maps out-of-range latents to blown-out colour.

Cost: CFG doubles inference compute — every step runs the network twice. Both branches are batched into one forward pass of batch size 2, so it costs roughly 2× the FLOPs and 2× the activation memory.

Negative prompts are the same mechanism. Pass the embedding of "blurry, watermark, extra fingers" as the unconditional branch instead of the empty one, and the formula pushes actively away from that text rather than merely away from nothing. They cost nothing extra: that forward pass was already happening.

Assembling the pipeline

All five components now have jobs: tokeniser and text encoder produce a 77×768 matrix; the U-Net predicts noise from a latent, a timestep and that matrix; the scheduler steps between timesteps; the VAE decoder converts the final latent to pixels.

Python
import torchfrom diffusers import StableDiffusionPipeline, DPMSolverMultistepSchedulerpipe = StableDiffusionPipeline.from_pretrained(    "stable-diffusion-v1-5/stable-diffusion-v1-5", torch_dtype=torch.float16).to("cuda")pipe.scheduler = DPMSolverMultistepScheduler.from_config(pipe.scheduler.config)image = pipe(    prompt="a red bicycle against a blue wall, golden hour",    negative_prompt="blurry, watermark, low quality",    num_inference_steps=25, guidance_scale=7.5,    generator=torch.Generator("cuda").manual_seed(42),).images[0]

The manual version below is worth reading once, because everything discussed above appears in it literally:

Python
import torch@torch.no_grad()def embed(text):    tok = pipe.tokenizer([text], padding="max_length", max_length=77,                         truncation=True, return_tensors="pt")    return pipe.text_encoder(tok.input_ids.to("cuda"))[0]embeddings = torch.cat([embed(""), embed("a red bicycle, blue wall")])latents = torch.randn((1, 4, 64, 64), device="cuda", dtype=torch.float16)pipe.scheduler.set_timesteps(25)latents = latents * pipe.scheduler.init_noise_sigmafor t in pipe.scheduler.timesteps:    inp = pipe.scheduler.scale_model_input(torch.cat([latents] * 2), t)    with torch.no_grad():        pred = pipe.unet(inp, t, encoder_hidden_states=embeddings).sample    eps_uncond, eps_cond = pred.chunk(2)                 # CFG    pred = eps_uncond + 7.5 * (eps_cond - eps_uncond)    latents = pipe.scheduler.step(pred, t, latents).prev_samplewith torch.no_grad():                                    # decode once, at the end    image = pipe.vae.decode(latents / pipe.vae.config.scaling_factor).sample

Schedulers: how you walk the chain matters

Training uses 1,000 timesteps; sampling does not have to. A scheduler picks which subset to visit and how to integrate between them. The reverse process is a differential equation and schedulers are numerical solvers for it — better solvers need fewer evaluations for the same accuracy.

SchedulerSteps for good outputDeterministic?CharacterUse when
DDPM~1000no (adds noise each step)the original; high diversityreference / research only
DDIM50–100yes (η = 0)reproducible; enables inversionimage editing, latent interpolation
Euler / Euler a20–30Euler yes, Euler-a nofast, forgiving, ancestral variant is creativegeneral purpose, quick iteration
DPM-Solver++ (2M)15–25yesbest quality-per-step availableproduction default
LCM / Turbo1–8no (re-noises between steps)distilled; needs a matching modelreal-time and interactive tools

Two traps. Swapping schedulers changes the image for a fixed seed, because different solvers follow different trajectories — a seed is not a portable identifier. And LCM and Turbo schedulers require distilled weights; pointing an LCM scheduler at stock SD 1.5 at 4 steps produces noise, and the failure looks like a scheduler bug rather than a model mismatch.

What this means when you build something

The architecture dictates where your problems will come from, and knowing the mechanism tells you which knob to reach for.

SymptomActual causeFix
Text and small faces come out mushyVAE compression discarded that detail at 64×64generate larger, or upscale then run img2img at low strength
Colours blown out, edges look friedguidance scale pushed latents outside the decoder's trained rangedrop guidance to 5–7, or enable guidance rescaling
End of a long prompt ignoredsilent truncation at 77 CLIP tokensshorten, front-load the important terms, or use a weighting library that chunks
Images look like grey noiseforgot the 0.18215 scale factor on encode or decodemultiply on encode, divide on decode
4 steps produces garbageLCM scheduler on non-distilled weightsload an LCM LoRA or a Turbo checkpoint, or raise steps to 20+
Out of memory at 1024×1024attention memory grows quadratically with latent areaenable attention slicing and VAE tiling (pipe.vae.enable_tiling()); use enable_model_cpu_offload()

On memory: float16 roughly halves weight memory and is near-free in quality for inference; enable_attention_slicing() computes attention in chunks, trading ~10% speed for a large saving; enable_model_cpu_offload() keeps only the currently-executing component on the GPU and gets SDXL into about 8 GB. That last one works precisely because the pipeline is five separable components running in sequence, not one monolith.

On cost: compute per step scales with latent area, so 512 to 1024 is 4× the latent pixels and more than 4× the attention cost. Halving step count by switching to DPM-Solver++ is usually a bigger and safer win than any other single change available to you.