Course Content
Generative AI System Design Interview
11 sections · 27 lessons
High-resolution synthesis: latent diffusion, training, fast serving and safety
The process from How diffusion works, slowly is correct and, run on pixels at high resolution, unaffordable. This lesson starts with the change that made image generation a product rather than a demonstration, then covers training, the serving techniques that cut the step count, and safety.
The problem, in numbers
A 512×512 RGB image is 512 × 512 × 3 = 786,432 values. A 1024×1024 image is 3,145,728.
Every one of the 50 network passes in How diffusion works, slowly operates on an array that size. The network's cost scales with the spatial area it processes, and any attention layers inside it scale worse than linearly in the number of spatial positions. Running 50 passes over three million values, per image, at 50 images per second, is not a product.
And here is the observation that unlocks it: most of those values are redundant. Neighbouring pixels in a photograph are highly correlated. There is far less information in an image than there are numbers describing it.
The fix
Do the diffusion somewhere smaller.
- Train an autoencoder that compresses an image into a much smaller array — the latent — and decodes it back into an image.
- Encode the entire training set into latents, once.
- Run the whole diffusion process in latent space. Add noise to latents, train the network to predict noise in latents, generate by denoising from random latent noise.
- Decode once at the very end to get the image.
With a typical 8× spatial downsampling factor, a 512×512×3 image becomes a 64×64×4 latent: 16,384 values instead of 786,432 — about 48 times smaller. The 50 network passes now run on the small array, and only the single final decode touches full resolution.
The measured compute reduction is large but not exactly 48×, because the network is redesigned to be relatively deeper and wider on the smaller array. Published latent diffusion work reports order-of-magnitude reductions in both training and inference cost, and that is the right way to state it.
The autoencoder, and why it is not a plain one
A plain autoencoder trained only to reconstruct produces blurry output, for exactly the reason given in the GANs and VAEs lesson: where several outputs are plausible, minimising reconstruction error favours their average.
Blurry decodes would cap the entire system's quality, so the autoencoder is trained with three objectives together:
- Reconstruction loss — get the pixels approximately right.
- Perceptual loss — match deep features, not pixels, so it preserves what looks right rather than what measures right.
- An adversarial loss — a discriminator judges the decoded output, which forces genuinely sharp texture instead of a plausible average.
The generative adversarial network from Section 7 has not disappeared. It is inside this, supplying sharpness.
The latent is also lightly regularised — kept close to a simple distribution, or quantised against a learned codebook — so that the diffusion model operates on a well-behaved space rather than an arbitrary one.
What is lost, and how to measure your ceiling
This is the part candidates miss, and it is the most practically useful idea in this section.
The autoencoder is a hard ceiling on everything downstream. Anything it cannot reconstruct can never be generated, no matter how good the diffusion model is. The diffusion model produces latents; the decoder turns latents into images; if the decoder cannot render something, it does not appear.
What typically suffers:
- Small text. Letters are high-frequency, structured, and unforgiving. This is a major contributor to the text-rendering failure that the text-to-image weaknesses lesson covers.
- Faces at small scale. A face occupying 40 pixels loses the detail that makes it coherent.
- Fine repeating patterns — fabric weave, brickwork, foliage — which decode as plausible texture rather than the specific pattern.
- Precise straight lines and grids, which acquire slight waviness.
The diagnostic, and run it before anything else: take a hundred real images of the kind you care about, encode and decode them, and look at the results side by side. Compute a perceptual distance (see the metrics lesson). Whatever is degraded there is your ceiling. If small text is unreadable after a round trip, no amount of diffusion training will make your model render small text, and you need a higher-capacity autoencoder or a lower downsampling factor.
The trade is direct: less compression means a better ceiling and more compute per step. A 4× downsample instead of 8× gives four times the latent area and roughly four times the per-step cost.
Training
Three things: the objective restated compactly, the noise schedule (which behaves differently at high resolution), and the guidance mechanism that Section 9 is entirely built on.
The objective
Sample an image, encode it to a latent, pick a random step t, add the corresponding noise, show the network the noisy latent and t, and train it to predict the noise. Average the squared error over many samples.
What it rewards, in one sentence: accurate noise prediction at every noise level. That is all. There is no adversary, no discriminator, no balance to maintain, and the loss curve descends monotonically enough to be informative — which is the main practical reason diffusion displaced adversarial training.
Noise schedules, and the high-resolution correction
The noise schedule decides how much noise is present at each step. A linear schedule adds noise at a constant rate; a cosine-shaped schedule spends more steps in the middle range where the interesting structural decisions are made, and generally produces better results.
There is a resolution-dependent subtlety worth knowing, because it explains a class of failure. Because neighbouring pixels are correlated, adding the same per-pixel noise destroys less information in a large image than in a small one — the redundancy means the structure survives longer. So a schedule tuned at 256×256 leaves too much signal intact at the noisy end when applied at 1024×1024, and the model is never trained on genuinely uninformative inputs.
The symptom is characteristic: generated images have decent local detail and a washed-out or implausible global composition, because the model never learned to make large-scale structural decisions from nothing. The fix is to shift the schedule towards more noise at higher resolutions. If you see that symptom, this is where to look.
Classifier-free guidance
The mechanism Section 9 depends on, and it is set up here, at training time, with a change small enough to state in one sentence.
At training time: drop the conditioning at random for some fraction of examples — 10% is typical — replacing it with a null token. The single network therefore learns two things at once: predicting noise given the conditioning, and predicting noise with no conditioning.
At sampling time: run the network twice per step, once conditioned and once unconditioned. Then take the difference between the two predictions — which is the part of the prediction the conditioning is responsible for — and exaggerate it by a factor called the guidance scale:
final prediction = unconditional + scale × (conditional − unconditional)
At scale 1 you get the ordinary conditional prediction. At scale 7 you get a prediction that has been pushed seven times further in the direction the conditioning pointed.
Three things follow, and they are the whole of Section 9's tuning story:
- Higher scale means stronger adherence to the conditioning, and reduced diversity and eventually visible artefacts — over-saturated colours, harsh contrast, over-simplified composition.
- It costs two network passes per step instead of one, so a 30-step generation with guidance is 60 passes. This doubles your inference bill and is often omitted from candidates' cost estimates.
- It requires no separate classifier, which is what the name is contrasting against and why it displaced the earlier approach.
Compute, stated plainly
Illustrative orders of magnitude, at 2026:
- Training a general-purpose latent diffusion base model from scratch: published figures for open models are in the region of 100,000+ accelerator-hours, which is roughly $250,000 at the $2.50 per hour from the cost and latency lesson, running for weeks on a substantial cluster.
- Fine-tuning an existing base model on a domain: hundreds to a few thousand accelerator-hours — $500 to $5,000.
- Training the autoencoder: a meaningful fraction of the base training cost, and it is usually taken pretrained rather than retrained.
The ratio is the point, and it is the decision table from The build, fine-tune, or prompt decision again. Almost nobody should be training a base diffusion model. Nearly everybody should be adapting one.
Serving: making a slow model fast enough
Latent diffusion made each step cheaper. The serving design takes fewer steps, and then prices the result.
Fewer steps, three ways
1. A better sampler. The reverse process is the numerical solution of a trajectory, and naive sampling takes small first-order steps. Higher-order solvers take larger, better-informed steps and reach comparable quality in far fewer of them — an illustrative move from 50 steps to around 20–25 at no visible cost. This requires no retraining: it is a generation-time choice, which makes it the first thing to try and the cheapest win available.
2. Step distillation. Train a student model to reproduce in a few steps what the teacher produces in many, by matching the teacher's multi-step trajectory. Distilled models reach usable quality at 1 to 4 steps, a 10–25× reduction against the baseline.
The honest limitations: quality at 1–2 steps is visibly below the many-step teacher on complex scenes, diversity narrows, and it requires a distillation training run per base model, so it does not come free. Where it wins is interactive products, where being able to regenerate in 200 ms changes what the product is.
3. Skip guidance where you can. Classifier-free guidance doubles the passes. Running it only on the earlier, structural steps and dropping it later recovers a meaningful fraction of the cost with a small quality change. Measure it; the crossover point is model-specific.
The quality-against-steps curve, which is the real decision
Read the shape, not the numbers. Quality rises steeply to around 8 steps, flattens by 25, and is indistinguishable past 30 — while cost keeps climbing linearly forever. Running 100 steps buys nothing and costs twice as much as 50.
The design decision is where on the flattening part to sit, and it should be a per-tier decision rather than a single global choice:
| Tier | Steps | Rationale |
|---|---|---|
| Live preview while the user types | 2–4, distilled | Latency is the product; quality is provisional |
| Standard generation | 20–25, good sampler | The knee of the curve |
| "High quality" on request | 40–50 | The user asked and will wait |
| Batch backfill | 25 | Nobody is waiting; cost is everything |
Offering a preview tier at all is a product insight worth stating: a 300 ms rough image that the user can accept or reject before committing to a two-second render cuts wasted full generations substantially, because most rejections happen on composition rather than on detail.
Batching and quantisation
Diffusion batches well: every image in the batch takes the same number of steps, so there is none of the ragged-length waste that continuous batching exists to solve. Batching 8 images gives close to 8× throughput at a modest latency cost. The limit is memory, since activations at high resolution are large.
Quantisation to 8-bit gives roughly the usual gains with a small quality cost that is more visible here than in text, because image artefacts are directly perceptible. Measure on your own outputs; do not assume.
Cost per image, computed
Illustrative, at $0.00069 per GPU-second, latent diffusion at 1024×1024, batch 8:
Standard tier — 25 steps with guidance (so 50 network passes):
- Denoising: 50 passes × 9 ms per pass per image at batch 8 = 0.45 GPU-s
- Autoencoder decode: 0.03 GPU-s
- Safety classification: 0.01 GPU-s
- Total: 0.49 GPU-s → $0.00034 per image
Preview tier — 4-step distilled, no guidance:
- Denoising: 4 passes × 9 ms = 0.036 GPU-s
- Decode: 0.03 GPU-s
- Total: 0.066 GPU-s → $0.000046 per image, roughly 7× cheaper
At 50 images per second sustained: 50 × 0.49 = 24.5 GPU-seconds of work per second, so about 25 accelerators for the standard tier at full utilisation, plus headroom. In money, 50 per second is 4.3 million images a day at $0.00034 = about $1,470 a day, or $536,000 a year.
Safety, monitoring and follow-ups
Four safety concerns specific to image generation, then the two extensions every interviewer asks about.
Output content filtering
Classify the generated image before returning it, with a model trained on the categories you must block. Roughly 10 ms and a fraction of a cent.
Two refinements worth naming, because they show you have thought about where diffusion differs:
- Filter early, not only at the end. At around step 8 of 25, the composition is already determined (the table in How diffusion works, slowly). Decoding a preview at that point and classifying it lets you abort the remaining 17 steps, saving roughly two thirds of the compute on rejected generations. Worth doing when the rejection rate is more than a few percent.
- Filter the prompt too, which Section 9 covers, because catching an intent before spending a GPU-second is cheaper than catching an image after.
Memorisation and the copyright exposure it creates
Diffusion models can reproduce training images near-verbatim. The effect is well documented, and the driver is well understood: images that appear many times in the training data are memorised far more strongly than images that appear once. A stock photograph reproduced across thousands of web pages is a strong memorisation candidate; an image appearing once is not.
That makes the primary mitigation a data-pipeline task, which is a satisfying and slightly counter-intuitive answer:
- Deduplicate the training data aggressively, including near-duplicates by perceptual hash and embedding similarity. This is the highest-value single control and it also improves training efficiency (see Data, licensing, and provenance).
- Check outputs against a training-image index. Embed generated images and look for very high similarity to a training image. Adds a vector lookup; catches the clearest cases.
- Log and monitor the rate of near-matches as a health metric.
- Track provenance and licences per training image, because when a question is asked you need to answer it with records rather than recollection.
Be accurate about the legal position: whether training on copyrighted images is permitted, and under what conditions, is contested and differs by jurisdiction, and it is actively changing. Keep the engineering claims — deduplication, output matching, provenance records — and leave the legal conclusions to people qualified to make them.
Watermarking and provenance
As Data, licensing, and provenance and the face generation safety lesson said: embed an invisible watermark in generated images and attach signed provenance credentials. Image watermarks survive re-compression and mild cropping reasonably well and are removable by a determined adversary; credentials are stripped by a screenshot. Both are good-faith labelling; neither is a control.
The two extensions
Upscaling as a separate stage. Generating natively at 2048×2048 is expensive — cost scales with area, so it is roughly four times a 1024×1024 generation. Generating at 1024 and then running a dedicated super-resolution model is substantially cheaper and is the standard architecture. The super-resolution model is itself usually a small diffusion or adversarial model conditioned on the low-resolution image, at a few steps.
The trade to state: upscaling adds detail that is plausible rather than derived from the original generation, so it can invent texture that the composition did not imply. For photographic subjects that is fine; for text and fine structure it makes things worse.
Inpainting and outpainting. Regenerate a masked region while keeping the rest fixed, or extend an image beyond its borders. The mechanism is a neat use of the process from How diffusion works, slowly: at every denoising step, take the known region, add the noise appropriate to the current step, and paste it over the corresponding part of the working image. The model therefore always sees correct context for the region it is generating, and the generated part is conditioned on it at every noise level.
This works reasonably with an unmodified model and works considerably better with a model fine-tuned on masked inputs, because a base model has never been trained on the sharp mask boundary and tends to produce seams.