Generative AI System Design Interview

Course Content

Generative AI System Design Interview

11 sections · 27 lessons

High-resolution synthesis: latent diffusion, training, fast serving and safety


The process from How diffusion works, slowly is correct and, run on pixels at high resolution, unaffordable. This lesson starts with the change that made image generation a product rather than a demonstration, then covers training, the serving techniques that cut the step count, and safety.

The problem, in numbers

A 512×512 RGB image is 512 × 512 × 3 = 786,432 values. A 1024×1024 image is 3,145,728.

Every one of the 50 network passes in How diffusion works, slowly operates on an array that size. The network's cost scales with the spatial area it processes, and any attention layers inside it scale worse than linearly in the number of spatial positions. Running 50 passes over three million values, per image, at 50 images per second, is not a product.

And here is the observation that unlocks it: most of those values are redundant. Neighbouring pixels in a photograph are highly correlated. There is far less information in an image than there are numbers describing it.

The fix

Do the diffusion somewhere smaller.

  1. Train an autoencoder that compresses an image into a much smaller array — the latent — and decodes it back into an image.
  2. Encode the entire training set into latents, once.
  3. Run the whole diffusion process in latent space. Add noise to latents, train the network to predict noise in latents, generate by denoising from random latent noise.
  4. Decode once at the very end to get the image.

With a typical 8× spatial downsampling factor, a 512×512×3 image becomes a 64×64×4 latent: 16,384 values instead of 786,432 — about 48 times smaller. The 50 network passes now run on the small array, and only the single final decode touches full resolution.

The measured compute reduction is large but not exactly 48×, because the network is redesigned to be relatively deeper and wider on the smaller array. Published latent diffusion work reports order-of-magnitude reductions in both training and inference cost, and that is the right way to state it.

The autoencoder, and why it is not a plain one

A plain autoencoder trained only to reconstruct produces blurry output, for exactly the reason given in the GANs and VAEs lesson: where several outputs are plausible, minimising reconstruction error favours their average.

Blurry decodes would cap the entire system's quality, so the autoencoder is trained with three objectives together:

  • Reconstruction loss — get the pixels approximately right.
  • Perceptual loss — match deep features, not pixels, so it preserves what looks right rather than what measures right.
  • An adversarial loss — a discriminator judges the decoded output, which forces genuinely sharp texture instead of a plausible average.

The generative adversarial network from Section 7 has not disappeared. It is inside this, supplying sharpness.

The latent is also lightly regularised — kept close to a simple distribution, or quantised against a learned codebook — so that the diffusion model operates on a well-behaved space rather than an arbitrary one.

What is lost, and how to measure your ceiling

This is the part candidates miss, and it is the most practically useful idea in this section.

The autoencoder is a hard ceiling on everything downstream. Anything it cannot reconstruct can never be generated, no matter how good the diffusion model is. The diffusion model produces latents; the decoder turns latents into images; if the decoder cannot render something, it does not appear.

What typically suffers:

  • Small text. Letters are high-frequency, structured, and unforgiving. This is a major contributor to the text-rendering failure that the text-to-image weaknesses lesson covers.
  • Faces at small scale. A face occupying 40 pixels loses the detail that makes it coherent.
  • Fine repeating patterns — fabric weave, brickwork, foliage — which decode as plausible texture rather than the specific pattern.
  • Precise straight lines and grids, which acquire slight waviness.

The diagnostic, and run it before anything else: take a hundred real images of the kind you care about, encode and decode them, and look at the results side by side. Compute a perceptual distance (see the metrics lesson). Whatever is degraded there is your ceiling. If small text is unreadable after a round trip, no amount of diffusion training will make your model render small text, and you need a higher-capacity autoencoder or a lower downsampling factor.

The trade is direct: less compression means a better ceiling and more compute per step. A 4× downsample instead of 8× gives four times the latent area and roughly four times the per-step cost.

Pixel diffusiondenoise directly in pixel space1024 × 1024 × 3 = 3,145,728 valuesevery one of the 50 steps operates on all of themLatent diffusionencode first, denoise in the latent space128 × 128 × 4 = 65,536 valuesabout 48× fewer values per stepImageEncoderLatentDenoise ×50Decoderthe expensive loop runs here, on the small tensorThe encoder and decoder run once each; the denoiser runs fifty times. Moving the loop into a 48× smaller space is the entire reason high-resolution image generation is affordable.
The saving compounds because it applies to every one of the fifty steps, not to the pipeline once.

Training

Three things: the objective restated compactly, the noise schedule (which behaves differently at high resolution), and the guidance mechanism that Section 9 is entirely built on.

Classifier-free guidance, per denoising stepConditionalpredictionUnconditionalpredictionExtrapolatethe gapTake thedenoise stepTwo forward passes per step, so guidance doubles the inference bill.
Turning guidance up sharpens prompt adherence and drains diversity and colour realism to pay for it.

The objective

Sample an image, encode it to a latent, pick a random step t, add the corresponding noise, show the network the noisy latent and t, and train it to predict the noise. Average the squared error over many samples.

What it rewards, in one sentence: accurate noise prediction at every noise level. That is all. There is no adversary, no discriminator, no balance to maintain, and the loss curve descends monotonically enough to be informative — which is the main practical reason diffusion displaced adversarial training.

Noise schedules, and the high-resolution correction

The noise schedule decides how much noise is present at each step. A linear schedule adds noise at a constant rate; a cosine-shaped schedule spends more steps in the middle range where the interesting structural decisions are made, and generally produces better results.

There is a resolution-dependent subtlety worth knowing, because it explains a class of failure. Because neighbouring pixels are correlated, adding the same per-pixel noise destroys less information in a large image than in a small one — the redundancy means the structure survives longer. So a schedule tuned at 256×256 leaves too much signal intact at the noisy end when applied at 1024×1024, and the model is never trained on genuinely uninformative inputs.

The symptom is characteristic: generated images have decent local detail and a washed-out or implausible global composition, because the model never learned to make large-scale structural decisions from nothing. The fix is to shift the schedule towards more noise at higher resolutions. If you see that symptom, this is where to look.

Classifier-free guidance

The mechanism Section 9 depends on, and it is set up here, at training time, with a change small enough to state in one sentence.

At training time: drop the conditioning at random for some fraction of examples — 10% is typical — replacing it with a null token. The single network therefore learns two things at once: predicting noise given the conditioning, and predicting noise with no conditioning.

At sampling time: run the network twice per step, once conditioned and once unconditioned. Then take the difference between the two predictions — which is the part of the prediction the conditioning is responsible for — and exaggerate it by a factor called the guidance scale:

final prediction = unconditional + scale × (conditional − unconditional)

At scale 1 you get the ordinary conditional prediction. At scale 7 you get a prediction that has been pushed seven times further in the direction the conditioning pointed.

Three things follow, and they are the whole of Section 9's tuning story:

  • Higher scale means stronger adherence to the conditioning, and reduced diversity and eventually visible artefacts — over-saturated colours, harsh contrast, over-simplified composition.
  • It costs two network passes per step instead of one, so a 30-step generation with guidance is 60 passes. This doubles your inference bill and is often omitted from candidates' cost estimates.
  • It requires no separate classifier, which is what the name is contrasting against and why it displaced the earlier approach.

Compute, stated plainly

Illustrative orders of magnitude, at 2026:

  • Training a general-purpose latent diffusion base model from scratch: published figures for open models are in the region of 100,000+ accelerator-hours, which is roughly $250,000 at the $2.50 per hour from the cost and latency lesson, running for weeks on a substantial cluster.
  • Fine-tuning an existing base model on a domain: hundreds to a few thousand accelerator-hours — $500 to $5,000.
  • Training the autoencoder: a meaningful fraction of the base training cost, and it is usually taken pretrained rather than retrained.

The ratio is the point, and it is the decision table from The build, fine-tune, or prompt decision again. Almost nobody should be training a base diffusion model. Nearly everybody should be adapting one.

Serving: making a slow model fast enough

Latent diffusion made each step cheaper. The serving design takes fewer steps, and then prices the result.

Fewer steps, three ways

1. A better sampler. The reverse process is the numerical solution of a trajectory, and naive sampling takes small first-order steps. Higher-order solvers take larger, better-informed steps and reach comparable quality in far fewer of them — an illustrative move from 50 steps to around 20–25 at no visible cost. This requires no retraining: it is a generation-time choice, which makes it the first thing to try and the cheapest win available.

2. Step distillation. Train a student model to reproduce in a few steps what the teacher produces in many, by matching the teacher's multi-step trajectory. Distilled models reach usable quality at 1 to 4 steps, a 10–25× reduction against the baseline.

The honest limitations: quality at 1–2 steps is visibly below the many-step teacher on complex scenes, diversity narrows, and it requires a distillation training run per base model, so it does not come free. Where it wins is interactive products, where being able to regenerate in 200 ms changes what the product is.

3. Skip guidance where you can. Classifier-free guidance doubles the passes. Running it only on the earlier, structural steps and dropping it later recovers a meaningful fraction of the cost with a small quality change. Measure it; the crossover point is model-specific.

The quality-against-steps curve, which is the real decision

Quality and cost against sampling steps (illustrative, latent diffusion at 1024x1024)10100125102050100log scalelog scaleQuality — % preference vs the50-step referenceCost — % of the 50-step cost
Quality and cost against sampling steps (illustrative, latent diffusion at 1024x1024)

Read the shape, not the numbers. Quality rises steeply to around 8 steps, flattens by 25, and is indistinguishable past 30 — while cost keeps climbing linearly forever. Running 100 steps buys nothing and costs twice as much as 50.

The design decision is where on the flattening part to sit, and it should be a per-tier decision rather than a single global choice:

TierStepsRationale
Live preview while the user types2–4, distilledLatency is the product; quality is provisional
Standard generation20–25, good samplerThe knee of the curve
"High quality" on request40–50The user asked and will wait
Batch backfill25Nobody is waiting; cost is everything

Offering a preview tier at all is a product insight worth stating: a 300 ms rough image that the user can accept or reject before committing to a two-second render cuts wasted full generations substantially, because most rejections happen on composition rather than on detail.

Batching and quantisation

Diffusion batches well: every image in the batch takes the same number of steps, so there is none of the ragged-length waste that continuous batching exists to solve. Batching 8 images gives close to 8× throughput at a modest latency cost. The limit is memory, since activations at high resolution are large.

Quantisation to 8-bit gives roughly the usual gains with a small quality cost that is more visible here than in text, because image artefacts are directly perceptible. Measure on your own outputs; do not assume.

Cost per image, computed

Illustrative, at $0.00069 per GPU-second, latent diffusion at 1024×1024, batch 8:

Standard tier — 25 steps with guidance (so 50 network passes):

  • Denoising: 50 passes × 9 ms per pass per image at batch 8 = 0.45 GPU-s
  • Autoencoder decode: 0.03 GPU-s
  • Safety classification: 0.01 GPU-s
  • Total: 0.49 GPU-s → $0.00034 per image

Preview tier — 4-step distilled, no guidance:

  • Denoising: 4 passes × 9 ms = 0.036 GPU-s
  • Decode: 0.03 GPU-s
  • Total: 0.066 GPU-s → $0.000046 per image, roughly 7× cheaper

At 50 images per second sustained: 50 × 0.49 = 24.5 GPU-seconds of work per second, so about 25 accelerators for the standard tier at full utilisation, plus headroom. In money, 50 per second is 4.3 million images a day at $0.00034 = about $1,470 a day, or $536,000 a year.

Safety, monitoring and follow-ups

Four safety concerns specific to image generation, then the two extensions every interviewer asks about.

Four controls around a pixel generatorDeduplicate training dataClassify prompt and outputWatermark invisiblyKeep a provenance record
Deduplication is the copyright control: images repeated in the corpus are the ones memorised verbatim.

Output content filtering

Classify the generated image before returning it, with a model trained on the categories you must block. Roughly 10 ms and a fraction of a cent.

Two refinements worth naming, because they show you have thought about where diffusion differs:

  • Filter early, not only at the end. At around step 8 of 25, the composition is already determined (the table in How diffusion works, slowly). Decoding a preview at that point and classifying it lets you abort the remaining 17 steps, saving roughly two thirds of the compute on rejected generations. Worth doing when the rejection rate is more than a few percent.
  • Filter the prompt too, which Section 9 covers, because catching an intent before spending a GPU-second is cheaper than catching an image after.

Memorisation and the copyright exposure it creates

Diffusion models can reproduce training images near-verbatim. The effect is well documented, and the driver is well understood: images that appear many times in the training data are memorised far more strongly than images that appear once. A stock photograph reproduced across thousands of web pages is a strong memorisation candidate; an image appearing once is not.

That makes the primary mitigation a data-pipeline task, which is a satisfying and slightly counter-intuitive answer:

  • Deduplicate the training data aggressively, including near-duplicates by perceptual hash and embedding similarity. This is the highest-value single control and it also improves training efficiency (see Data, licensing, and provenance).
  • Check outputs against a training-image index. Embed generated images and look for very high similarity to a training image. Adds a vector lookup; catches the clearest cases.
  • Log and monitor the rate of near-matches as a health metric.
  • Track provenance and licences per training image, because when a question is asked you need to answer it with records rather than recollection.

Be accurate about the legal position: whether training on copyrighted images is permitted, and under what conditions, is contested and differs by jurisdiction, and it is actively changing. Keep the engineering claims — deduplication, output matching, provenance records — and leave the legal conclusions to people qualified to make them.

Watermarking and provenance

As Data, licensing, and provenance and the face generation safety lesson said: embed an invisible watermark in generated images and attach signed provenance credentials. Image watermarks survive re-compression and mild cropping reasonably well and are removable by a determined adversary; credentials are stripped by a screenshot. Both are good-faith labelling; neither is a control.

The two extensions

Upscaling as a separate stage. Generating natively at 2048×2048 is expensive — cost scales with area, so it is roughly four times a 1024×1024 generation. Generating at 1024 and then running a dedicated super-resolution model is substantially cheaper and is the standard architecture. The super-resolution model is itself usually a small diffusion or adversarial model conditioned on the low-resolution image, at a few steps.

The trade to state: upscaling adds detail that is plausible rather than derived from the original generation, so it can invent texture that the composition did not imply. For photographic subjects that is fine; for text and fine structure it makes things worse.

Inpainting and outpainting. Regenerate a masked region while keeping the rest fixed, or extend an image beyond its borders. The mechanism is a neat use of the process from How diffusion works, slowly: at every denoising step, take the known region, add the noise appropriate to the current step, and paste it over the corresponding part of the working image. The model therefore always sees correct context for the region it is generating, and the generated part is conditioned on it at every noise level.

This works reasonably with an unmodified model and works considerably better with a model fine-tuned on masked inputs, because a base model has never been trained on the sharp mask boundary and tends to produce seams.