Course Content
Generative AI System Design Interview
11 sections · 27 lessons
Personalized headshots: framing, identity metrics and adaptation techniques
Design a product where a user uploads a handful of photographs of themselves and receives professional-looking headshots of that person in settings and outfits they did not photograph.
This is the most practical section in the image half of the course, because it is what a product team is actually asked to build. It is also the section where the interview question is as much about unit economics and consent as about models.
Clarifying questions
- How many input photos may we require? Every extra photo improves identity fidelity and costs conversion. Twenty photos is a better model and a worse funnel. Assume 10–15, which is a common compromise.
- What turnaround is acceptable? Seconds, minutes, or "we will email you"? This is the question that decides the architecture, because per-user training cannot be done in seconds. Assume 20–30 minutes, delivered by notification.
- How many outputs? Assume 60 images across 6 styles, from which the user picks a few.
- What identity fidelity is needed? "Recognisably you to a stranger" and "recognisably you to your mother" are different bars, and the second is much harder.
- What is the price point? Ask this. It is not a product-manager question in this design; it determines how many GPU-minutes you may spend per user, which determines the adaptation technique compared below.
- What volume, and how bursty? Assume 5,000 users a day with heavy peaks, which Serving economics shows is the number that actually decides viability.
The framing
Subject-driven adaptation of a pretrained generator.
Three words worth unpacking:
- Subject-driven — the system must learn one specific person from a few examples, then render that person in new contexts. This is different from Section 9, where the prompt described a generic subject.
- Adaptation — not training. The decision from The build, fine-tune, or prompt decision lands squarely on the cheap end here, and it is not close: nobody pretrains a diffusion model per user.
- Pretrained generator — everything from Sections 8 and 9 is assumed. The base model already knows what a person, an office, and studio lighting look like. You are teaching it one new proper noun.
The three things in tension, named upfront
The whole case study is a negotiation between three properties, and stating the triangle early is a strong opening:
| Property | What it means | What pushing it costs |
|---|---|---|
| Identity fidelity | It looks like this specific person | Pushed too far, outputs become copies of the input photos |
| Flexibility | It can put them in new poses, outfits, and settings | Requires the adaptation not to dominate the base model |
| Cost per user | GPU-minutes and storage | The cheapest techniques have the lowest fidelity |
You cannot have all three at maximum. Which corner you sit in is a product decision driven by the price point, which is why you asked about it.
Metrics
Three measurable properties, one human judgement, and a specific trap in who does the judging.
Identity similarity
The core metric, and it has a clean definition.
Take a face-recognition model — the kind trained to say whether two photographs show the same person — and use it as an embedding function. Embed each of the user's input photographs and average them into a reference embedding. Embed the face in each generated image. The identity score is the cosine similarity between the two.
The number is meaningless without calibration, so calibrate it. On a held-out set of real photographs, measure the similarity distribution for pairs of photographs of the same person and for pairs of different people. Those two distributions give you a threshold and, importantly, a ceiling: two genuine photographs of the same person taken years apart may score lower than you expect, and a generated image cannot be held to a stricter standard than reality.
State the threshold in those terms — "above the 20th percentile of genuine same-person pairs" — rather than as a bare number, which will not transfer between face models.
Image quality and prompt adherence
Quality as in the high-resolution metrics lesson, plus a domain-specific check: headshots have specific failure modes — waxy skin, distorted ears, asymmetric glasses, hands where hands should not be — and a small trained artefact detector on those pays for itself.
Prompt adherence as in the text-to-image metrics lesson. Did you get the requested outfit, background, and lighting? This is the metric that catches over-adaptation, below.
The tension, made concrete
Push identity fidelity hard — more training steps, higher adaptation strength — and something specific happens:
- Identity similarity rises. Good.
- Prompt adherence falls. The outputs start reproducing the settings, outfits, and even the poses of the input photographs, because the adaptation has learned the whole distribution of those photos rather than the person in them.
- Diversity collapses. Sixty outputs start to look like sixty crops of the same four images.
At the extreme, the system returns the user's own photographs back to them with slight variation, which scores perfectly on identity and is worthless as a product.
So identity similarity must always be reported with prompt adherence and output diversity. A rise in one with a fall in the others is over-adaptation, and it is the characteristic failure of this product. Measure diversity as the average pairwise distance between output embeddings.
The human judgement, and who should make it
"Does this look like me" decides whether the user asks for a refund, and no automatic metric predicts it reliably.
Adaptation techniques compared
The heart of the case study. Five techniques, and the decision table is the deliverable.
The comparison
Illustrative figures for a latent diffusion base model at 1024×1024, at $2.50 per GPU-hour.
| Full fine-tune | Subject-token | Textual inversion | LoRA | Encoder-based | |
|---|---|---|---|---|---|
| Per-user training | 45 GPU-min | 30 GPU-min | 8 GPU-min | 12 GPU-min | none |
| Compute cost per user | $1.88 | $1.25 | $0.33 | $0.50 | ~$0.00 |
| Storage per user | ~6 GB | ~4 GB | ~5 KB | ~60 MB | none |
| Identity fidelity | Highest | Very high | Moderate | High | Moderate |
| Flexibility retained | Lowest — forgets | High with prior preservation | High | High | High |
| Time to first image | ~50 min | ~35 min | ~12 min | ~16 min | seconds |
| One-off engineering | Low | Low | Low | Low | Very high |
| Best for | Nothing, at this scale | Quality-first, low volume | Cheap experiments | The default | Very high volume, instant results |
Read the LoRA column. 27% of the cost, 1% of the storage, 94% of the quality. That ratio is why it is the default answer, and being able to state it in that form is the point of this comparison.
The recommendation, and when to deviate
Default to LoRA. It sits at the knee of every curve here, the adapters are small enough to store per user indefinitely and to swap in milliseconds at serving time, and it retains the base model's flexibility because the base weights are untouched.
Deviate in two directions:
- Towards subject-token fine-tuning if identity fidelity is the product and volume is low enough that $1.25 per user is affordable — a premium tier, or a business-to-business offering. Use prior preservation, or you will find every subsequent generation drifting towards your most recent subject.
- Towards encoder-based if instant results are the product or if volume is very high. The economics invert completely: enormous one-off engineering and training investment, then essentially free per user. The catch is fidelity — a single reference photo carries less information than fifteen, and it shows — and the catch behind the catch is that the encoder is a research project, not a fortnight.
A pragmatic architecture worth proposing: encoder-based for an instant preview, LoRA for the final deliverable. The user sees themselves in a new setting within seconds, which is what sells the product, while the good images train in the background. That answer combines the two and is usually better than either.