Course Content
Generative AI System Design Interview
11 sections · 27 lessons
Text-to-image: serving, safety and control extensions
The diffusion serving lesson priced diffusion. This lesson adds what conditioning changes: guidance doubles the passes, retries multiply requests, and the user is waiting. It then covers the harms text-to-image concentrates in one feature, and the control extensions interviewers ask about next.
The latency budget
Target: something visible in under a second, finished image in under four.
| Stage | Latency |
|---|---|
| Prompt safety classification | 20 ms |
| Text encoding (both encoders) | 15 ms |
| Denoising, 25 steps × 2 passes for guidance, batch 4 | 2,300 ms |
| Autoencoder decode | 90 ms |
| Output safety classification | 25 ms |
| Watermark embedding and provenance metadata | 15 ms |
| Transfer | 120 ms |
| Total | ~2,585 ms |
Denoising is 89% of it, which tells you where every optimisation belongs.
Progressive preview
The most valuable product feature in this lesson, and it falls out of the observation in How diffusion works, slowly that composition is settled early.
Decode the intermediate latent at steps 4, 8, and 16 and stream those images to the user. At step 4 the user sees a rough composition — about 400 ms in. At step 8 they can tell whether it is the right idea. At step 16 it is nearly final.
Three benefits, and the third is the one to say in an interview:
- Perceived latency collapses from 2.6 seconds to about 400 ms, exactly as streaming did for text in the chatbot serving lesson.
- The user can cancel early, which saves the remaining steps on generations they were going to reject anyway.
- It converts Section 8's early-abort safety check into a user-facing feature — you were going to decode an early preview for content classification regardless, so the marginal cost of showing it is one small decode.
The cost is real: each preview decode is roughly 90 ms of autoencoder work. Three previews add about 270 ms of compute to a 2,300 ms generation — around 12%. Worth it.
Queueing, rate limiting, and cost control
Image generation is expensive per request and easy to abuse, so admission control is a first-class component, not an operational afterthought.
- A queue with a visible position and estimate. People wait for images if they can see the wait.
- Per-user rate limits, on requests per minute and on generations per day. Tie them to the account tier.
- A cost ceiling per user per period, enforced in generations rather than in currency so it is understandable.
- Priority tiers so paying users are not queued behind bulk free traffic.
- Degrade under load by dropping the step count — from 25 to 16 — before dropping requests. A slightly softer image beats an error, and most users will not notice.
- Reject early when the queue exceeds the timeout.
Two more controls specific to this system. Cache by (prompt, seed, model version, parameters) — exact repeats are common, particularly from bots and from users refreshing. And fix and return the seed, so a user who liked an image can reproduce it and vary it, which is both a genuine feature for designers and a way to avoid re-rolling generations at random.
Cost per delivered image, computed
Illustrative, at $0.00069 per GPU-second, 1024×1024, batch 4, 25 steps with guidance:
- Prompt safety: 0.005 GPU-s
- Text encoding: 0.01 GPU-s
- Denoising, 50 passes: 0.60 GPU-s
- Three preview decodes plus the final decode: 0.12 GPU-s
- Output safety: 0.01 GPU-s
- Total: 0.745 GPU-s → $0.00051 per generation
Now the correction most candidates miss: users do not keep the first image. With an illustrative 2.4 generations per image the user actually keeps, the cost per delivered image is $0.00123.
At 20 generations per second sustained: 20 × 0.745 = 14.9 GPU-seconds per second, so roughly 15 accelerators at full utilisation plus headroom, and 1.7 million generations a day at $0.00051 = about $880 a day, or $320,000 a year.
Two levers change that materially. Early cancellation via preview cuts abandoned generations — saving perhaps 15% at an illustrative cancellation rate. And a distilled 4-step preview tier (see the diffusion serving lesson) for exploratory generations, with full quality only on a final render, can cut the bill by more than half for users who iterate.
Safety, monitoring and follow-ups
Text-to-image concentrates several distinct harms in one feature, and the controls differ by harm. Treat them separately.
Prompt filtering and output classification
Prompt filtering runs first because it is cheapest — 20 ms and a fraction of a cent to avoid spending a GPU-second. Classify the prompt against your policy categories and block or rewrite.
Its limits are the same as every input filter (see Safety as architecture), and are worth stating specifically here: paraphrase defeats keyword matching; describing a prohibited thing without naming it defeats concept matching; prompts in other languages are frequently under-covered by filters trained mainly on English; and encoded, transliterated, or deliberately misspelled prompts slip through.
Output classification is therefore mandatory, not optional. Classify the generated image before returning it. Add the early-step check from the high-resolution safety lesson to abort expensive generations that were heading somewhere prohibited.
Both together still miss things, so keep the human review sample and the report path.
Generating recognisable real people
The harm from Section 7 arrives here through a text interface, which makes it easier.
Controls, and their gaps:
- Name blocking in prompts. Maintain a list of public figures and block prompts naming them. Cheap, and defeated by description — "the current president of...", or a detailed physical description without a name.
- Face-embedding checks on the output. Embed any generated face and compare against an index of public figures, discarding above a similarity threshold. Catches what name blocking misses, costs about 10 ms, and covers only people in your index.
- Refuse photorealism plus a named individual as a combined condition, allowing clearly stylised depictions if your policy does.
Be clear that these together are a strong deterrent, not a guarantee, and that private individuals — who are not in any index — are covered only by the fact that the model was not trained specifically on them.
Style imitation and artist consent
"In the style of [living artist]" is the most genuinely unsettled question in this case study, and the honest answer distinguishes candidates.
The engineering controls available:
- Name-based blocking for artists who have requested it, plus an opt-out request process with a public record. This is the control most systems implement first.
- Training-data opt-out honoured at crawl time and at dataset-construction time, with the documented exclusion process from the data lesson.
- Style-similarity detection on outputs, comparing against a registry of representative works. Technically possible and considerably less reliable than identity matching, because style is diffuse where a face is specific.
- Attribution and provenance so a generated image is labelled as such.
And the honest framing: whether training on published artwork is permissible, and whether style is protectable at all, differs by jurisdiction, is contested, and is actively being litigated and legislated. State the engineering controls with confidence; do not state a legal conclusion. An interviewer asking this question is testing judgement, not legal knowledge.
Watermarking and provenance
As before: invisible watermark plus signed provenance credentials on every generated image, understood as good-faith labelling rather than a control. Both are removable; both are worth doing.
Monitoring
Track the prompt-block rate and output-block rate as a pair — a divergence means the prompt filter has been worked around. Track the alignment score and the compositional prompt suite from the metrics lesson on every model change. Track generations per delivered image, which rises when quality or alignment drops and is an excellent early warning that costs nothing to collect. And sample outputs for human review, weighted towards prompts that scored near a safety threshold rather than uniformly.
The control extensions
These are the standard follow-up questions:
- Image-to-image. Start the denoising process from a noised version of a supplied image rather than from pure noise, at an intermediate step. How far back you noise it is the "strength" dial: noise it lightly and the output stays close to the input; noise it heavily and only the composition survives.
- Structural conditioning — depth maps, edge maps, human pose skeletons. An auxiliary network processes the control input and injects features into the denoiser alongside the text conditioning, giving spatial control that text alone cannot express. This is the practical answer to the spatial-relation weakness from The known weaknesses for users who can supply a layout.
- Negative prompts. Replace the unconditional branch of classifier-free guidance with a branch conditioned on text you want to avoid. The guidance computation then pushes away from that description as well as towards the prompt. It costs nothing extra, because that second pass was already being computed.
- Regional prompting, where different areas of the image get different prompts, applied by masking the cross-attention.