Generative AI System Design Interview

Course Content

Generative AI System Design Interview

11 sections · 27 lessons

Personalized headshots: framing, identity metrics and adaptation techniques


Design a product where a user uploads a handful of photographs of themselves and receives professional-looking headshots of that person in settings and outfits they did not photograph.

This is the most practical section in the image half of the course, because it is what a product team is actually asked to build. It is also the section where the interview question is as much about unit economics and consent as about models.

The product, from upload to deliveryUseruploads12 photosVerify itis themTrain asmall adapterGenerate100 imagesDeliverthe best 20Adapter training happens once per user and dominates the per-user cost.
Identity fidelity, image quality and unit cost pull against each other, and the adapter is where all three meet.

Clarifying questions

  • How many input photos may we require? Every extra photo improves identity fidelity and costs conversion. Twenty photos is a better model and a worse funnel. Assume 10–15, which is a common compromise.
  • What turnaround is acceptable? Seconds, minutes, or "we will email you"? This is the question that decides the architecture, because per-user training cannot be done in seconds. Assume 20–30 minutes, delivered by notification.
  • How many outputs? Assume 60 images across 6 styles, from which the user picks a few.
  • What identity fidelity is needed? "Recognisably you to a stranger" and "recognisably you to your mother" are different bars, and the second is much harder.
  • What is the price point? Ask this. It is not a product-manager question in this design; it determines how many GPU-minutes you may spend per user, which determines the adaptation technique compared below.
  • What volume, and how bursty? Assume 5,000 users a day with heavy peaks, which Serving economics shows is the number that actually decides viability.

The framing

Subject-driven adaptation of a pretrained generator.

Three words worth unpacking:

  • Subject-driven — the system must learn one specific person from a few examples, then render that person in new contexts. This is different from Section 9, where the prompt described a generic subject.
  • Adaptation — not training. The decision from The build, fine-tune, or prompt decision lands squarely on the cheap end here, and it is not close: nobody pretrains a diffusion model per user.
  • Pretrained generator — everything from Sections 8 and 9 is assumed. The base model already knows what a person, an office, and studio lighting look like. You are teaching it one new proper noun.

The three things in tension, named upfront

The whole case study is a negotiation between three properties, and stating the triangle early is a strong opening:

PropertyWhat it meansWhat pushing it costs
Identity fidelityIt looks like this specific personPushed too far, outputs become copies of the input photos
FlexibilityIt can put them in new poses, outfits, and settingsRequires the adaptation not to dominate the base model
Cost per userGPU-minutes and storageThe cheapest techniques have the lowest fidelity

You cannot have all three at maximum. Which corner you sit in is a product decision driven by the price point, which is why you asked about it.

Metrics

Three measurable properties, one human judgement, and a specific trap in who does the judging.

The tension you have to measure, not resolvePush identity similarity• Face embedding distance to the uploads• More adapter steps, higher similarity• Output startsreproducing the input photosPush quality and adherence• Prompt adherence to outfit and setting• Fewer steps keeps the base model's range• Output stops looking like the user
The person judging must be the user themselves — strangers cannot tell whether a face is the right one.

Identity similarity

The core metric, and it has a clean definition.

Take a face-recognition model — the kind trained to say whether two photographs show the same person — and use it as an embedding function. Embed each of the user's input photographs and average them into a reference embedding. Embed the face in each generated image. The identity score is the cosine similarity between the two.

The number is meaningless without calibration, so calibrate it. On a held-out set of real photographs, measure the similarity distribution for pairs of photographs of the same person and for pairs of different people. Those two distributions give you a threshold and, importantly, a ceiling: two genuine photographs of the same person taken years apart may score lower than you expect, and a generated image cannot be held to a stricter standard than reality.

State the threshold in those terms — "above the 20th percentile of genuine same-person pairs" — rather than as a bare number, which will not transfer between face models.

Image quality and prompt adherence

Quality as in the high-resolution metrics lesson, plus a domain-specific check: headshots have specific failure modes — waxy skin, distorted ears, asymmetric glasses, hands where hands should not be — and a small trained artefact detector on those pays for itself.

Prompt adherence as in the text-to-image metrics lesson. Did you get the requested outfit, background, and lighting? This is the metric that catches over-adaptation, below.

The tension, made concrete

Push identity fidelity hard — more training steps, higher adaptation strength — and something specific happens:

  1. Identity similarity rises. Good.
  2. Prompt adherence falls. The outputs start reproducing the settings, outfits, and even the poses of the input photographs, because the adaptation has learned the whole distribution of those photos rather than the person in them.
  3. Diversity collapses. Sixty outputs start to look like sixty crops of the same four images.

At the extreme, the system returns the user's own photographs back to them with slight variation, which scores perfectly on identity and is worthless as a product.

So identity similarity must always be reported with prompt adherence and output diversity. A rise in one with a fall in the others is over-adaptation, and it is the characteristic failure of this product. Measure diversity as the average pairwise distance between output embeddings.

The human judgement, and who should make it

"Does this look like me" decides whether the user asks for a refund, and no automatic metric predicts it reliably.

Adaptation techniques compared

The heart of the case study. Five techniques, and the decision table is the deliverable.

The comparison

Illustrative figures for a latent diffusion base model at 1024×1024, at $2.50 per GPU-hour.

Full fine-tuneSubject-tokenTextual inversionLoRAEncoder-based
Per-user training45 GPU-min30 GPU-min8 GPU-min12 GPU-minnone
Compute cost per user$1.88$1.25$0.33$0.50~$0.00
Storage per user~6 GB~4 GB~5 KB~60 MBnone
Identity fidelityHighestVery highModerateHighModerate
Flexibility retainedLowest — forgetsHigh with prior preservationHighHighHigh
Time to first image~50 min~35 min~12 min~16 minseconds
One-off engineeringLowLowLowLowVery high
Best forNothing, at this scaleQuality-first, low volumeCheap experimentsThe defaultVery high volume, instant results
Adaptation techniques, indexed against full fine-tuning = 100 (illustrative)125102050100Encoder-basedTextual inversionLoRASubject-tokenFull fine-tunelog scaleCompute cost per user (% of fullfine-tune)Storage per user (% of full fine-tune)Identity fidelity (% of full fine-tune)
Adaptation techniques, indexed against full fine-tuning = 100 (illustrative)

Read the LoRA column. 27% of the cost, 1% of the storage, 94% of the quality. That ratio is why it is the default answer, and being able to state it in that form is the point of this comparison.

The recommendation, and when to deviate

Default to LoRA. It sits at the knee of every curve here, the adapters are small enough to store per user indefinitely and to swap in milliseconds at serving time, and it retains the base model's flexibility because the base weights are untouched.

Deviate in two directions:

  • Towards subject-token fine-tuning if identity fidelity is the product and volume is low enough that $1.25 per user is affordable — a premium tier, or a business-to-business offering. Use prior preservation, or you will find every subsequent generation drifting towards your most recent subject.
  • Towards encoder-based if instant results are the product or if volume is very high. The economics invert completely: enormous one-off engineering and training investment, then essentially free per user. The catch is fidelity — a single reference photo carries less information than fifteen, and it shows — and the catch behind the catch is that the encoder is a research project, not a fortnight.

A pragmatic architecture worth proposing: encoder-based for an instant preview, LoRA for the final deliverable. The user sees themselves in a new setting within seconds, which is what sells the product, while the good images train in the background. That answer combines the two and is usually better than either.