Generative AI System Design Interview

Course Content

Generative AI System Design Interview

11 sections · 27 lessons

Text-to-image: framing, alignment metrics and re-captioned data


Design a system that turns a written description into an image.

Section 8 built a machine that produces good images. This section makes it produce the image someone asked for, and that is where nearly all the remaining difficulty lives.

What "aligned to the prompt" actually demandsNamed objects presentCounts correctAttributes bound correctlySpatial relations right
Systems clear the top rung long before the bottom ones, which is why alignment must be scored separately.

Clarifying questions

  • General-purpose or a specific style? A model that renders one house style — product photography, a game's art direction, technical illustration — is a fine-tune of a base model and a much easier problem than open-ended generation.
  • What resolution and aspect ratios? Fixed square output is simpler; arbitrary aspect ratios require training-time and serving-time handling (see the serving lesson).
  • What latency? Interactive with a preview, or a background job with a notification? Assume interactive.
  • How much control does the user get? A prompt only, or also negative prompts, reference images, pose and depth control, seeds, and per-region control? Each is a separate component.
  • Who are the users? Consumers who type a sentence and designers who iterate twenty times on one image want different products, and the second group's needs — reproducibility, fine control, variation from a fixed seed — are more demanding.
  • Volume. Assume 20 images per second at peak.

The framing

Conditional generation, where the difficulty is alignment rather than quality.

Be precise about the split, because it structures the rest of the answer:

  • Image quality — is this a good photograph or illustration? Largely inherited from Section 8. Diffusion at scale produces beautiful images.
  • Prompt alignment — is it the image that was asked for? Not inherited from anywhere. It has to be designed for, in the data (the data step below), in the conditioning architecture (Conditioning the diffusion process), and in the guidance scale, and it remains partly unsolved (The known weaknesses).

A candidate who spends the round on quality has designed Section 8 again. Say the split out loud in the first two minutes and then spend your time on the second half.

What "aligned" actually requires

The prompt "a red cube on top of a blue sphere, on a wooden table, in soft morning light" requires the system to get five different things right:

RequirementDifficulty
The right objects present — cube, sphere, tableReliable
The right global style — soft morning lightReliable
The right attributes bound to the right objects — red cube, blue sphereUnreliable
The right count — one of eachUnreliable
The right spatial relation — cube on top of sphereUnreliable

The first two are solved. The last three are the open problems of the field and are covered in The known weaknesses. Knowing which is which — and saying so — is the single clearest signal of whether a candidate has actually worked with these systems.

Metrics

Two axes that must be measured separately, because a system can move up one while falling down the other and a combined score would hide it.

Two axes that must not be averagedAxis 1 — image quality• FID against a reference distribution• Artefacts, hands, faces, symmetry• Improves with more sampling stepsAxis 2 — prompt alignment• CLIP score, and its ceiling effects• Question answering over the image• Improves with stronger guidance
Guidance moves one axis up and the other down, so a single combined score hides the trade entirely.

Axis 1: image quality

Fréchet Inception Distance, with every caveat from the high-resolution metrics lesson — the downsampling blindness at high resolution, the sample-size sensitivity, and the fact that a preference-tuned model scores worse while being preferred. Supplement with patch-level FID and human pairwise preference.

Axis 2: prompt alignment

The cheap instrument: image-text similarity. Score the generated image against the prompt using a contrastive model of the kind the image captioning training lesson described — the same machinery that supplied the text encoder in Conditioning the diffusion process.

It is fast, needs no reference, and runs on live traffic. Its weakness is exactly the failure class you most need to detect. Contrastive image-text models behave substantially like bags of words: they are strong on which concepts are present and weak on how those concepts relate. On "a red cube on a blue sphere" they score an image with a blue cube on a red sphere almost as highly, because all four concepts are present.

So the cheap metric is blind to attribute binding, counting, and spatial relations — the three things that are actually broken.

The better instrument: question-based evaluation. Decompose the prompt into simple checkable questions, then ask a vision-language model (see the image captioning model lesson) each one about the generated image:

PromptGenerated questions
"a red cube on top of a blue sphere, on a wooden table"Is there a cube? Is the cube red? Is there a sphere? Is the sphere blue? Is the cube above the sphere? Is there a wooden table?

Score the fraction answered correctly. This measures precisely what similarity scoring misses, and it degrades gracefully — a partial answer produces a partial score, which is far more useful for tracking progress than a single opaque number.

Two limitations to state: the evaluating vision-language model has its own weaknesses on counting and spatial relations, so it will be wrong sometimes and its errors correlate with the generator's (the shared-blind-spot problem from Evaluating output with no correct answer); and generating good questions from a prompt is itself a model call that can be wrong. Validate a sample against human judgement, exactly as you would calibrate any judge.

Human evaluation, on both axes separately

Show a rater the prompt and two images. Ask two questions:

  1. Which image is better as an image?
  2. Which image better matches the prompt?

Ask them separately, and report them separately. Combining them into "which do you prefer" is the most common mistake in this evaluation, because raters weight beauty heavily and you will systematically ship models that make prettier images that ignore instructions.

Keep a fixed compositional prompt suite — several hundred prompts deliberately loaded with counts, spatial relations, attribute bindings, negations, and rendered text — and track alignment on it on every release. This suite will look embarrassing for a long time. That is what makes it useful.

Data

The insight of the data step is short and consequential: caption quality, not image quality, is the binding constraint on prompt-following.

Re-captioning is the whole fixRaw web alt-textFilter byCLIP scoreRe-captionwith a VLMMixsyntheticand rawTrainKeep some original captions so the model still understands how people actually write prompts.
Prompt-following is bounded by caption quality, not image quality — the pixels were always fine.

What the raw data looks like

Billions of image-text pairs scraped from the web, filtered as in the image captioning data lesson. Enormous, free, and — for this purpose — badly shaped.

Consider what web alt-text actually contains for a photograph of a red cube on a blue sphere:

red cube blue sphere 3d render stock photo

That is typical. Short, keyword-ordered, no grammar, no relations. Now consider what a user types:

"a red cube balanced on top of a blue sphere, sitting on a wooden table in soft morning light"

A model trained on the first learns to respond to keyword lists. It learns that "red", "cube", "blue", and "sphere" should all appear somewhere in the image. It has never been trained on a caption that says on top of, so it has no reason to have learned what that means visually.

That is the mechanism behind the failures in The known weaknesses. The architecture is not incapable of representing spatial relations. It was never given data in which spatial relations were described.

Re-captioning, which is the fix

Run a vision-language captioning model (see the image captioning model lesson) over your training images and generate long, detailed, grammatical captions describing objects, attributes, counts, spatial relationships, style, and lighting. Train on those instead of, or alongside, the originals.

This is a well-documented and highly effective technique, and it is the answer an interviewer is listening for when they ask how you would improve prompt-following.

It has three real costs, and naming them is what makes the answer credible:

  • Compute. Captioning a billion images at the $0.00006 each from the image captioning serving lesson is roughly $60,000 and a substantial batch job. Real, and affordable relative to training.
  • Inherited blind spots. The captioner's weaknesses become the generator's. If the captioner cannot count, no caption says "three", and the generator will not learn to count. This is a genuine ceiling, and it is why re-captioning improves relations more than it improves counting.
  • Distribution lock-in. Train only on synthetic captions and the model responds well to captioner-style prose and worse to how real users write — which is terse, ungrammatical, and keyword-heavy. The standard mitigation is a mixture: train on a blend of original and synthetic captions, so the model handles both registers. The blend ratio is a real tuning parameter.

The rest of the filtering pipeline

  • Image-text similarity threshold to drop mismatched pairs.
  • Aesthetic scoring — a small model trained on human ratings, used to weight or filter towards better-looking images. This is a large part of why modern outputs look good, and it is also where a homogenised house style creeps in, because you are training towards one rating model's taste.
  • Resolution and aspect-ratio filters, plus aspect-ratio bucketing so training batches contain images of one shape and non-square outputs work at serving time.
  • Watermark and overlay detection, or the model learns to draw stock-photo watermarks.
  • Deduplication, which as the high-resolution safety lesson explained is the primary control on memorisation.
  • Safety filtering of categories you will not generate.

On consent and provenance

Everything in Data, licensing, and provenance applies with more force here, because the training images are creative works with identifiable authors. Store provenance and licence per image, honour opt-out signals at crawl time and maintain a documented exclusion process, and keep records good enough to answer a question about a specific image. The safety lesson returns to the style question.