Generative AI System Design Interview

Course Content

Generative AI System Design Interview

11 sections · 27 lessons

High-resolution synthesis: framing, metrics and how diffusion works


Design a system that generates high-resolution photographic images.

This is the technical core of the image half of the course. Sections 9, 10, and 11 are all diffusion case studies and all assume this one. The walk-through of how diffusion works, later in this lesson, is the explanation to read twice.

The two levers for reaching 1024 pixelsCascade of upsamplers• Generate 64px, then 256, then 1024• Each stage is a separate model to train• Errors from stage one are amplifiedLatent diffusion• Diffuse in a compressed latent space• One model, decoder restores pixels• Autoencoder caps the achievable detail
Both avoid diffusing directly in pixel space, where attention cost scales with the pixel count squared.

Clarifying questions

  • What resolution? 512×512, 1024×1024, and 2048×2048 differ by factors of four in pixel count and considerably more in cost. Assume 1024×1024 as the delivered output.
  • What throughput? Images per second at peak decides the whole capacity plan. Assume 50 images per second at peak.
  • Interactive or batch? A user waiting in a browser needs something on screen in a few seconds. A background job for a catalogue does not. Assume interactive, which is the harder case.
  • Conditioned on what? This section keeps it unconditional or lightly conditioned. Section 9 adds text prompts, which is where most of the difficulty moves.
  • What is the quality bar? "Indistinguishable from a photograph" and "good enough for a thumbnail" are separated by most of the cost.
  • Is a fixed style acceptable? A narrow style is much easier and much cheaper than general-purpose photorealism.

The framing

Iterative denoising: start from random noise and refine it towards an image over many passes through a network.

The consequence arrives immediately and dominates the section. Every other system in this course runs its network once per output unit — one pass per token, one pass per caption word. Diffusion runs it N times per image, where N is somewhere between 4 and 100.

So the design problem here is not "what architecture". It is:

How do we get N down, and what does each reduction cost in quality?

That question is answered by the serving design, and everything before it in this section exists to make it answerable.

The two levers, named upfront

There are exactly two ways to make diffusion cheaper, and everything in this section is one of them:

  1. Make each step cheaper. Latent diffusion does this, by roughly an order of magnitude. Quantisation adds a little more.
  2. Take fewer steps. Better samplers and step distillation (see the serving lesson) do this, from 50 steps to as few as 1–4.

They compose. Together they are the difference between a research demonstration and a product.

Metrics

The face generation metrics lesson introduced Fréchet Inception Distance and its blind spots. Two things change at high resolution, and both are easy to get wrong.

Which instrument sees which failureyesnearlyblindcheapyesyescheappartlynocheapyesyesslowCompositionFine detailCostFIDLPIPSCLIP scoreHuman A/BFID resizes images to 299 pixels before scoring them.
The metric is downsampling away exactly the high-frequency detail that high resolution was for.

FID is nearly blind to high-frequency detail

The feature extractor FID uses operates at a small fixed input size — a few hundred pixels on a side. Every 1024×1024 image is therefore downsampled before it is scored.

Whatever distinguishes a beautifully sharp 1024×1024 image from a soft one is largely destroyed by that downsampling. FID can barely see it. Two models whose outputs differ visibly in crispness, texture, and fine detail can post nearly identical FID scores.

This is the single most important measurement fact in this section, and it has a direct consequence: FID cannot be your quality gate at high resolution. It remains a useful check that you are producing the right kind of image, and it stops being a measure of how good the image is.

Two supplements to name:

  • Patch FID. Crop random patches at native resolution and compute FID on the patches. The patches are small enough not to be downsampled away, so high-frequency detail is preserved. This is the cheapest fix and it belongs in every high-resolution evaluation.
  • Sharpness and artefact detectors. Cheap classical measures of high-frequency energy, plus a trained detector for the specific artefacts your model produces. Not sophisticated; they catch the regressions FID sleeps through.

Perceptual similarity metrics

Perceptual distance metrics — which compare two images by the distance between their deep features rather than pixel by pixel — align far better with human judgement than pixel error. They need a reference image, so they measure reconstruction rather than generation.

That makes them the right tool for one specific and important job in this section: measuring what the autoencoder in the latent diffusion lesson loses. Encode a real image, decode it, and compare the result with the original. That number is a hard ceiling on your generated quality, and the latent diffusion design explains why.

Human preference, which dominates

For high-resolution generation, human pairwise preference is not a supplement to the automatic metrics; it is the metric, with the automatics acting as tripwires.

Run it as in the chatbot metrics lesson: same prompt or same seed, two models, both orderings shown, tie option, confidence intervals reported. For images, add a second question beyond "which is better" — "which has visible defects, and where?" — because it produces actionable output rather than a score.

How diffusion works, slowly

The most important explanation in the second half of this course. Diffusion models, explained plainly gave the outline; this part walks it with an image in mind at every stage and adds the pieces the case studies need.

The forward process: destroy an image on purpose

Take a photograph. Add a small amount of random noise to every pixel. Repeat, many times.

Concretely, with a 1,000-step schedule, here is what you would see:

StepWhat the image looks like
0The original photograph. A lighthouse on a cliff at dusk.
100Identical at a glance. Slight grain in the sky.
300Visibly noisy, like a high-ISO photograph. Every object still identifiable.
500Heavy static. You can make out a dark vertical mass against a lighter field.
700Shapes are gone. Faint large-scale brightness variation remains.
900Almost pure static. A very slight overall colour cast survives.
1000Pure random noise. Nothing of the photograph remains.

Two properties of this process do all the work.

It has no learned parameters. It is a fixed recipe — at step t, add exactly this much noise. You can run it on any image, at any step, instantly, for free.

It manufactures perfect training labels. For any image and any step you know both the noisy version and precisely what noise you added. That is a supervised regression problem with an exact, free label, which is why diffusion training is stable where the adversarial training in Section 7 is not.

The network, and what it predicts

The network sees a noisy image and the step number t, and predicts the noise that was added.

Two details matter for reasoning about the system.

Its shape. The classic architecture is a U-shaped network: a downsampling path that reduces spatial resolution while increasing channels, a bottleneck, an upsampling path back to full resolution, and skip connections carrying features across from each downsampling level to the matching upsampling level.

The shape fits the task precisely. The bottleneck sees the whole image at low resolution and can reason about global structure — where the horizon is, where the subject sits. The skip connections carry fine, high-frequency detail directly across, so it does not have to survive the squeeze through the bottleneck. Denoising needs both global understanding and local detail preservation, and this shape supplies both. Many recent large models replace the U-shaped network with a transformer operating on patches of the image (a diffusion transformer); what goes in and what comes out are unchanged, and so is everything else in this walk-through.

It is told the noise level. The step number t is embedded and injected into every layer. Without it, the network would not know whether it is looking at a lightly-grained image needing a small correction or near-pure static needing a structural guess. One network handles all noise levels because it is told which one it is at.

The reverse process: generating

Start with pure random noise — a 1024×1024 grid of random values containing no information at all.

For each of N steps, from the noisiest to the cleanest:

  1. Ask the network: given this image and this noise level, what noise is present?
  2. Use that prediction to compute an estimate of the clean image.
  3. Step partway towards it — not all the way — landing at the noise level of the next step.
  4. Add a small amount of fresh random noise, appropriate to that next level.

Here is what you would see at each stage of a 50-step run:

StepWhat the image looks like
0 (start)Pure static.
5Large soft blobs of colour. A darker region has appeared on one side.
12A recognisable composition: a vertical dark mass, a horizon line, sky above.
20Identifiably a tower on a cliff. Edges soft, texture absent.
30Stone texture appears. The lamp housing resolves. Sky gains gradient.
40Fine detail: individual stones, the railing, spray at the cliff base.
50A sharp finished image.

Structure first, then shapes, then texture. Nobody programmed that ordering. It falls out of the noise schedule: at high noise almost nothing survives, so the only thing recoverable is large-scale layout, and by the time detail is recoverable the layout is already fixed.

t = 1000t = 800t = 600t = 400t = 250t = 100t = 0Each step removes a little of the predicted noiseNo step produces a picture; each one produces a slightly less noisy tensor. The image only exists at the end.the cost dialEvery step is a full forward pass through the denoiser. 50 steps is 50 forward passes per image — which is why step count,not model size, is usually the first thing tuned in production.
Seven frames of a fifty-step walk — the shape only becomes recognisable in the last third.

The analogy, and three punctures

Diffusion is often introduced as sculpting a statue from marble. Take the opening and then correct it, because each misconception below produces wrong reasoning in an interview.

Puncture 1 — nothing is hidden in the marble. There is no image inside the noise waiting to be uncovered. The same starting noise with a different model, or a different prompt, gives a completely different picture. The information comes from the weights and the conditioning; the noise supplies only which of the many possible images you land on.

Puncture 2 — it is not pure removal. Most samplers add fresh noise back at every step. The trajectory is a stochastic walk towards the data distribution, not a monotone subtraction. Removing noise deterministically all the way, without adding any back, is a different sampler with different behaviour — usually less diverse output.

Puncture 3 — "revealing" is the wrong verb. At step 5 the network is not revealing anything. It is making a guess at what clean image this noise is consistent with, and that guess is extremely vague. Each subsequent step makes a less vague guess. The image is being decided progressively, not uncovered.

Sampler and model are different things

The last idea here, and the one the serving design depends on entirely.

The model is the trained network that predicts noise. It is fixed after training.

The sampler is the algorithm that decides how many steps to take, where to place them across the noise range, and how to move from one to the next. It is a numerical solver, and it is chosen at generation time.

That separation is why step counts can be cut dramatically without retraining. A better solver reaches the same destination in fewer, larger steps. You can switch samplers on a deployed model this afternoon and halve your cost, which is unusual enough to be worth saying out loud.