Generative AI System Design Interview

Course Content

Generative AI System Design Interview

11 sections · 27 lessons

Realistic faces: framing, image metrics and face data


Design a system that generates photorealistic images of human faces that do not belong to any real person.

This is the generative adversarial network section, and it is the right place to teach generative evaluation, because faces are the one image domain where every reader is an expert. A slightly wrong ear, a mismatched pair of earrings, or teeth that do not line up are invisible to a metric and instantly visible to a person.

Ask what it is for before designingUses that justify building• Balancing a biased training set• Stock avatars with no consent burden• Film and game background crowdsUses that forbid shipping• Fake identities and account fraud• Non-consensual imagery of real people• Fabricated evidence or news photos
The same generator serves both lists, so the purpose has to be settled before the architecture.

Clarifying questions

  • Unconditional, or controlled by attributes? Pure sampling — press a button, get a face — is a very different system from one where a user asks for a face with particular characteristics. Assume unconditional for the core design, with Control and editing adding control.
  • What resolution? 256×256 and 1024×1024 differ by a factor of 16 in pixels and far more in training difficulty.
  • What throughput? One face on demand, or ten million as a batch job?
  • Must the faces be verifiably not real people? A generated face that closely resembles a living person is a problem whatever your intent, and checking for it is a real component (see Safety, monitoring and follow-ups).
  • And first: what is this for?

Ask what it is for, before you design anything

This question is not a courtesy. Three plausible answers give three different systems:

Use caseWhat changes
Synthetic training data for a face-detection or blurring modelDemographic balance becomes the primary quality metric, not realism. Diversity and coverage matter more than any single image being beautiful. Disclosure is internal.
User avatarsStyle consistency and controllability dominate. Users must not be able to generate a face resembling a specific real person. Provenance labelling is a product requirement.
Stock imagery of people who do not existPhotorealism is the product. Licensing and disclosure obligations are highest, and the "does this resemble a real person" check becomes a hard gate.

There is also an answer that should stop the design: "to make images of a specific person". That is Section 10's problem with consent, or it is a deepfake. Establishing this in the first two minutes is what a responsible engineer does, and interviewers notice when a candidate designs a face generator for thirty minutes without asking who it is for.

The framing

Unconditional image generation: sample a random vector, produce an image drawn from the distribution of real face photographs.

"Unconditional" means there is no prompt, no attribute list, and no reference photograph. The only input is randomness. Everything the model produces comes from the training distribution and the particular random vector, which has an important consequence for the data step below: the output distribution is the training distribution. If the data is skewed, the output is skewed, and there is no prompt to correct it with.

Why a GAN, when Section 8 is about diffusion

Two honest reasons, and a candidate who gives them is not being nostalgic:

  • Inference speed. A generative adversarial network produces an image in one forward pass — around 20 ms. Diffusion needs 25 to 50 passes — around one second. That is roughly a 50× difference, and for a service generating millions of faces, or generating them inside an interactive loop, it is decisive.
  • The problem is narrow. Aligned, cropped faces are one of the most constrained image distributions there is. Adversarial training struggles on broad, diverse distributions and does well on narrow ones, and this is about as narrow as image generation gets.

If the interviewer asks why not diffusion, the answer is cost and latency, and the honest addition is that a diffusion model would likely produce somewhat better and more diverse faces if you could afford the inference.

Metrics for images

You cannot compute accuracy on a generated image. What you can do is compare the distribution of generated images against the distribution of real ones, and then look at individual images with human eyes. Both are necessary and neither is sufficient.

What a distribution distance cannot seeFID cannot seeSingle-image qualityMemorised samplesFine texture detailSmall-sample biasWhat humans dislike
FID compares two distributions, so it is structurally incapable of judging the image in front of you.

FID's blind spots, all of which matter

  • It depends entirely on the feature extractor. The classifier was trained on a particular dataset, so FID measures distance in that model's notion of perceptual similarity, which is better tuned to objects than to faces.
  • It is not comparable across sample sizes. FID computed on 5,000 samples and on 50,000 samples are different numbers for the same model, and the bias is systematic. Fix the sample count — 50,000 is the usual convention — and report it every time. Comparing two FIDs computed differently is a common and invalidating error.
  • It is sensitive to preprocessing. Resizing method, JPEG compression, and colour handling all move the number. Fix them and record them.
  • It says nothing about any individual image. A model can score well while producing occasional grotesque failures, because a small number of outliers barely move a distributional statistic.
  • Memorisation scores beautifully. A model that reproduces its training images has a nearly perfect FID and zero value. This is not hypothetical, and it is why a duplicate-detection check against the training set belongs in your evaluation.
  • It conflates quality with coverage. A model producing excellent images of half the distribution and a model producing mediocre images of all of it can score the same.

That last point has a standard fix worth naming: precision and recall for generative models, which separate the two. Precision asks what fraction of generated samples fall within the real distribution — a fidelity measure. Recall asks what fraction of the real distribution is covered by generated samples — a diversity measure. Mode collapse (see the training lesson) shows up as high precision with collapsed recall, which a single FID will not tell you.

What human raters catch that no metric does

Run two protocols.

Real-versus-fake discrimination. Show raters single images and ask whether each is a photograph or generated. The metric is the rate at which raters are fooled, and 50% means indistinguishable. It is intuitive, it is directly meaningful, and it degrades as raters learn the tells — so refresh your rater pool.

Pairwise preference between models, exactly as in the chatbot metrics lesson, with both orderings shown.

And the specific defects raters find that distributional metrics never will, all of them characteristic of face generation:

  • Asymmetry that should not be there — mismatched earrings, one eye a different shape or colour, uneven glasses frames.
  • Teeth, which are numerous, regular, and unforgiving.
  • Hands, when they appear.
  • Text in the background — signs, logos, clothing prints — which comes out as letter-shaped nonsense.
  • Background incoherence — a wall that changes material halfway across, or a shoulder that merges into scenery.
  • Hair boundaries and fine strands, which is where blending artefacts live.

None of these moves FID appreciably, and every one of them makes an image read as fake at a glance.

Data

Because generation is unconditional, the data question and the output question are the same question. Whatever is in the training set is what the model will produce.

The training mix is the output mix62 percent71 percent4.114 percent9 percent9.86 percent3 percent12.418 percent11 percent10.2Train shareOutput shareGroup FIDLighter skinDarker skinAge over 60Non-WesternIllustrative figures; the pattern is the point.
Under-represented groups come out rarer than they went in, because the generator chases the dense modes.

Where face data comes from, and the consent problem

Research face datasets are typically assembled by collecting photographs from public photo-sharing sites, often filtered by licence, then cropping and aligning the faces. Several widely-used datasets have been withdrawn or restricted after concerns were raised about whether the people in them had meaningfully consented, and that pattern should be treated as the norm rather than an exception when planning.

The engineering consequences, which is what an interview is asking about:

  • Store provenance per image — source, licence, collection date, and any consent record. If a removal request arrives and you cannot identify which images are affected, you cannot comply.
  • Plan for withdrawal. Assume some fraction of your training data will need removing during the model's life. That means a retraining schedule, and it is a further argument for keeping training runs affordable.
  • Licence terms travel. A licence that permits research use may not permit a commercial product. Check before training, not after.
  • Deletion from weights is not available. As Data, licensing, and provenance said, you can remove an image from a corpus and you cannot remove it from a trained model. The practical answer is scheduled retraining and being honest about the limitation.

Alignment, and what it costs you

Nearly every high-quality face generator is trained on aligned faces: detect facial landmarks, then rotate, scale, and crop so the eyes sit at fixed pixel positions in every image.

This helps enormously. It removes pose and scale variation from the distribution, so the model spends its capacity on faces rather than on framing, and it is a large part of why face generation reached photorealism before general image generation did.

It also constrains the output. Every generated face is centred, front-facing, and cropped the same way. That is fine for avatars and wrong for synthetic detection training data, where you need varied poses, occlusions, and distances — the whole point of that dataset. Alignment is a trade decided by the use case from the framing above.

It also leaves a fingerprint: because the eyes land in almost identical positions in every output, overlaying many generated faces shows the eyes aligned to the pixel. That has been a practical detection heuristic for aligned-face generators, and it is worth knowing both as a defender and as a limitation of the technique.

Demographic balance, which is the whole quality story