Course Content
Generative AI System Design Interview
11 sections · 27 lessons
Personalized headshots: the per-user pipeline, unit economics and consent
Per-user training is a scheduling and cost problem before it is a machine-learning problem. This lesson builds the pipeline, then works out what it costs and what it owes the people whose faces it handles.
Stage 1 — upload and validation
The stage that decides your refund rate, and the one that gets skipped in interview answers. Validate before you spend a GPU-minute:
- Count and resolution. Reject below a minimum; warn below a recommended count.
- Is there exactly one clear face per photo? Reject group shots, or crop with confirmation.
- Is it the same person across the set? Embed every face and cluster. Outliers are a photo of someone else, and one contaminating photo degrades the whole adapter. This check is cheap and it is the highest-value validation in the list.
- Variety. Fifteen photographs from one session in one outfit at one angle produce an adapter that has learned the outfit. Measure spread in pose and lighting and prompt the user for more if it is low.
- Quality. Reject heavy blur, heavy filters, very low resolution, and screenshots.
- Consent and identity — covered in the safety step below, and it belongs here in the flow.
Return every rejection with a specific reason and a way to fix it. A user who uploads again is worth far more than a refund.
Stage 2 — preprocessing
Detect and crop faces with margin, align if your technique benefits from it, resize to the training resolution, and generate captions binding the subject token. Generate the prior-preservation set if the technique needs one.
Stage 3 — the training job
A queued, bounded job on a worker pool. Twelve GPU-minutes for a low-rank adapter.
Design points that matter:
- Bounded worker pool, not a GPU per user. Concurrency is capped by your accelerator fleet, and the queue is where demand above that cap waits.
- Checkpoint and make it restartable. A twelve-minute job is short enough to run on interruptible capacity, which the economics below show is worth a great deal.
- Hyperparameters by input-set size. Ten photos and thirty photos want different step counts; a fixed recipe over-fits the small sets.
- Validate at the end of training. Generate four probe images with fixed prompts, score identity and diversity, and fail the job before delivery if it over-adapted. Catching it here costs one retry; catching it at the user costs a refund and a reputation.
Stage 4 — adapter storage
Write the adapter to object storage keyed by user, with the base model version, the technique, the hyperparameters, and a retention expiry recorded alongside. Sixty megabytes per user is trivially cheap; the expiry is the part that matters, for the reasons in the safety step below.
Stage 5 — generation
Load the base model, apply the adapter, and generate the batch: 60 images across 6 prompt styles, with fixed seeds recorded so any image can be reproduced or varied.
The serving trick that makes this affordable: the base model stays resident in accelerator memory and only the small adapter is swapped, in milliseconds. So one worker can serve many users' generations back to back without reloading gigabytes. This is the single strongest operational argument for low-rank adaptation over full fine-tuning, and it is easy to miss.
Stage 6 — post-processing and quality gating
- Face restoration and upscaling (see the high-resolution safety and follow-ups lesson).
- Safety classification on every output.
- Identity scoring on every output, discarding those below threshold.
- Aesthetic and artefact scoring, ranking the survivors.
- Deliver the best 40 of 60, ordered.
Generating more than you deliver is deliberate. Per-image generation is cheap relative to the training job, and the identity gate has a real failure rate, so over-generating and filtering is the cheapest quality control available.
Stage 7 — delivery and cleanup
Notify, deliver, and start the retention clocks: input photographs deleted after training, outputs after the download window, adapters at expiry.
Serving economics
This is the step that decides whether the product exists. Work it fully; it is what the interviewer wants to see.
The compute cost per user
Illustrative, at $2.50 per GPU-hour ($0.00069 per GPU-second):
| Item | Work | Cost |
|---|---|---|
| Validation (face detection, clustering) | 5 GPU-s | $0.003 |
| Adapter training, 12 GPU-min | 720 GPU-s | $0.497 |
| Generation, 60 images at 0.5 GPU-s | 30 GPU-s | $0.021 |
| Post-processing: upscale, restore, score, classify | 12 GPU-s | $0.008 |
| Storage, 60 MB adapter for 90 days plus outputs | $0.002 | |
| Subtotal | 767 GPU-s | $0.531 |
| Retries at an illustrative 10% failure rate | $0.053 | |
| Total compute | $0.584 |
At a $19 price point that looks like a 97% gross margin, and if you stop there you have given a naive answer.
The number that actually decides viability: utilisation
Accelerators are rented by the hour whether they are working or not. The cost above assumes 100% utilisation — every GPU-second you pay for is a GPU-second of user work.
Real demand for a consumer product is bursty. Traffic follows a daily cycle, a weekly cycle, and marketing spikes. If you provision for peak and average 25% utilisation, your effective cost per user is four times the figure above: $2.34, not $0.58.
That is the number to compute in the interview, and it changes the answer:
| Utilisation | Effective compute cost per user | Gross margin at $19 |
|---|---|---|
| 100% | $0.58 | 97% |
| 50% | $1.17 | 94% |
| 25% | $2.34 | 88% |
| 10% | $5.84 | 69% |
Still viable at $19. Now run it at a $5 price point with 10% utilisation and the margin is negative before you have paid for payment processing, support, or engineers. Price point and utilisation together decide the technique in Adaptation techniques compared, which is why the framing lesson told you to ask.
The four levers on utilisation
- Interruptible capacity. Training jobs are 12 minutes and checkpointed, which makes them close to ideal for pre-emptible instances. Illustratively 60–80% cheaper than on-demand. This is the largest single lever and it is available because of a design choice made in the per-user pipeline above.
- A cheap overnight tier. Offer a lower price for delivery within 12 hours and use it to fill the troughs. Users self-select, and your utilisation curve flattens without any capacity change.
- Batch the queue into windows rather than starting each job on arrival. Slightly worse latency, materially better packing.
- Share the fleet. Run lower-priority batch work — re-captioning, evaluation runs, re-generation of expired outputs — on idle capacity.
The rest of the unit economics
Compute is not the whole cost, and a strong answer says so:
- Payment processing — around 3% of revenue, so about $0.57 on $19.
- Refunds. "It doesn't look like me" is the dominant reason, and at an illustrative 8% refund rate that is $1.52 per sale in lost revenue plus the compute already spent. The identity gate in the per-user pipeline above is a revenue control, not a quality nicety — it is the cheapest way to move this number.
- Support, which correlates with refunds.
- Storage and egress, small but real at 40 delivered images per user.
Take $19 revenue, $2.34 effective compute at 25% utilisation, $0.57 processing, $1.52 refunds, and $0.40 support and storage: about $14.17 contribution per sale, roughly 75%. That is the answer to give — a real margin with the terms named, not a 97% figure that ignores how accelerators are billed.
Adapter lifecycle
Adapters are small, and keeping them forever is still wrong.
- Hot cache of recently-used adapters near the generation fleet, so a returning user generating more images does not wait.
- Cold storage for the rest.
- Expiry at a defined retention period. This aligns two things that usually conflict: it is cheaper, and it is required. An adapter trained on someone's face is derived biometric data (see the safety step below), and deleting it on a schedule is the behaviour you want by default.
- Regeneration on demand if an expired user returns — 12 GPU-minutes, and honest to charge for.
Safety, monitoring and follow-ups
This product takes photographs of people's faces, derives a model of their likeness, and generates new images of them. Every one of those three steps carries an obligation.
Consent verification: does this face belong to this user?
The core control, because everything else follows from it. Without it you have built a service for generating images of anyone whose photographs an attacker can find.
The strongest available control is a live capture. During onboarding, capture a short video or a guided sequence of selfies in the application, with a liveness check to defeat a held-up photograph. Then verify that the uploaded training photographs match the live capture using the same face-embedding comparison as the metrics lesson. Uploads that do not match the live capture are rejected.
Blocking non-consensual and celebrity generation
Layer three checks:
- On upload, compare the input faces against a public-figure index and reject matches. Someone uploading photographs of a well-known person is not personalising their own headshots.
- On upload, the liveness match above, which is the general case of the same control.
- On output, score generated faces against the public-figure index as in the text-to-image safety lesson, which catches the case where an adapter drifts towards a well-known face.
State the gap honestly: none of this protects a private individual whose photographs someone else possesses, beyond the liveness check. The liveness check is therefore not one control among several — it is the one that matters, and weakening it for conversion is the decision that turns this product into a harm.
Biometric data, retention, and deletion
Face photographs, face embeddings, and an adapter trained on someone's face are all data about a person's body. Several jurisdictions treat information of this kind as a special category with stricter requirements around consent, purpose limitation, and deletion, and the details differ and are changing.
Design to the strict end and you will be close to right in most places:
- Retention by default, not on request. Input photographs deleted after training. Outputs deleted after the download window. Adapters deleted at expiry. Live captures deleted immediately.
- A deletion path that is tested. Deletion must reach object storage, caches, the hot adapter cache, backups, logs, and any analytics copy. Run a periodic test that deletes a synthetic user and verifies every store. An untested deletion path is a policy, not a capability.
- Deletion from a trained model is not possible. This is a further argument for adapters: the subject exists in a 60 MB file that can be deleted, not in base weights that cannot.
- Data residency, if you operate across regions.
Quality failures and the refund path
The output does not look like the user. Detect it before they do:
- Identity score per output, with delivery blocked below threshold.
- Automatic retry with adjusted hyperparameters — more or fewer steps, different adaptation strength — when the batch scores low overall. One retry costs $0.50 and saves a refund worth $19.
- Ask for more photographs when the input set was the cause, which the validation metrics from the per-user pipeline above usually reveal.
- A clear refund path, because some fraction will fail regardless and a fought refund costs more than it saves.
Monitoring
Track the identity score distribution over time, since a base-model update can shift it silently. Track the refund rate by input-photo count, which tells you whether your minimum is set right. Track queue depth and wait times, which are the product experience. Track utilisation, which is the margin. And track the upload rejection rate by reason, because a spike in "not the same person" is either a bad instruction screen or an attack.
The extensions
- Products instead of people. The same machinery, with the face-embedding identity metric replaced by a general image-embedding similarity. Consent becomes straightforward, which makes it commercially attractive: a merchant's product photographed once and rendered in any setting.
- Pets. Popular, technically identical, and free of the biometric obligations.
- Multiple subjects in one image — two people who both need to be recognisable. Substantially harder, because it hits the attribute-binding problem from The known weaknesses with the highest possible stakes: the two identities blend.
- Video headshots, which is Section 11's problem with an identity constraint on every frame.