Course Content
Generative AI System Design Interview
11 sections · 27 lessons
Realistic faces: the style-based GAN, training, editing and deepfake safeguards
GANs, VAEs, and why diffusion mostly displaced them introduced the adversarial game. This lesson makes it concrete, adds the two architectural ideas that took face generation to photorealism, and then follows the design through training, editing, and the safeguards a face generator cannot ship without.
The game
The generator takes a random vector — typically 512 numbers drawn from a standard normal distribution, called z — and produces an image. It never sees a real photograph.
The discriminator takes an image and outputs a score for how likely it is to be real. It sees both real images and generator output.
They train in alternation:
- Draw real images and generate fake ones. Train the discriminator to score reals high and fakes low.
- Generate fresh fakes. Train the generator to make the discriminator score them high — with the discriminator's weights held fixed, so the gradient flows back through it into the generator.
- Repeat, millions of times.
The generator learns only through the discriminator. It has no direct access to a real image and no pixel-level target. Every piece of information it receives about what a face looks like comes through a single scalar critique — which is why the process is powerful and why it is fragile.
Progressive growing
Training directly at 1024×1024 is unstable: the discriminator can trivially distinguish real from fake early on, because early generator output is noise at every scale, and the generator receives a useless signal.
Progressive growing starts both networks at 4×4 pixels, trains until stable, then fades in a new layer to double the resolution, and repeats: 8×8, 16×16, and so on to 1024×1024. Each stage starts from a working model at half the resolution and only has to learn the added detail.
Two benefits: it stabilises training, and it is substantially faster, because most of the training happens at low resolution where each step is cheap.
The style-based architecture
The idea that made faces controllable, and the reason the control and editing later in this lesson is possible.
In a plain generator, the random vector z is fed in at the input and everything downstream is derived from it. The style-based design changes two things:
- A mapping network transforms z through several fully-connected layers into an intermediate vector w. The generator's actual input is a learned constant, not z.
- w is injected at every resolution level, controlling the normalisation statistics of the features at that level — a mechanism usually called adaptive instance normalisation. Separate per-pixel random noise is added at each level for stochastic detail such as individual hair strands and skin pores.
Why this matters: different resolution levels control different scales of the image. Injecting a different w at the coarse levels changes pose, face shape, and hair length. At the middle levels it changes facial features and hairstyle. At the fine levels it changes colour and micro-texture. That separation is what makes attribute editing possible.
The mapping network also matters for a subtler reason. z comes from a fixed spherical distribution, but real face attributes are not distributed spherically — some combinations are common and some are rare or absent. Forcing a spherical latent onto that reality entangles attributes. The mapping network can warp the space, and the resulting w space is measurably less entangled, which is what control and editing depends on.
Training
Adversarial training is the least reliable training process in this course. Knowing why, and knowing the standard stabilisers, is the practical content of this step.
Why it is unstable
Two networks optimise against each other, so neither has a fixed target. Three consequences:
- The loss values do not tell you how it is going. A healthy-looking pair of loss curves is compatible with excellent output and with noise. You evaluate by looking at samples and by computing FID on a schedule.
- One network can overpower the other. A discriminator that becomes too strong provides gradients the generator cannot act on; a weak one provides no useful signal at all. The balance has to be maintained through the whole run.
- It can diverge late. A run that has looked good for days can collapse, which is why you snapshot frequently and keep the best checkpoint by FID rather than the last one.
Mode collapse
The characteristic failure. The generator discovers a small set of outputs the discriminator scores highly and produces only those, with minor variation.
The insidious part is that the images can be excellent. The generator has found something that genuinely fools the discriminator; it has stopped covering the distribution. Nothing in the loss objects, because nothing in the loss asks about coverage.
Detect it with the generative recall metric from Metrics for images, and by generating a large grid of samples and looking. On faces it is unmistakable once you see it — pages of the same person with different lighting.
The stabilisers that are standard practice
- A non-saturating generator loss, which keeps gradients useful early in training when the discriminator is winning easily.
- A gradient penalty on the discriminator, or spectral normalisation of its weights, which limits how sharply it can change its verdict and is one of the largest single stability wins available.
- Different learning rates for the two networks, so the balance can be tuned directly.
- An exponential moving average of the generator's weights for sampling. The averaged generator produces noticeably better images than the raw one at any given step, and it costs nothing but a copy of the weights. Do this; it is the cheapest quality gain in training.
- Discriminator augmentation when data is limited. Randomly augment the images the discriminator sees — both real and fake, with the same pipeline — so it cannot win by memorising the training set. This is what makes training on tens of thousands of images rather than millions feasible.
- Frequent snapshots and best-checkpoint selection by FID, because of the late-divergence risk.
Compute, stated honestly
Illustrative for a 1024×1024 style-based face generator, at 2026 orders of magnitude: on the order of 1,000 to 5,000 accelerator-hours, or $2,500 to $12,500 at the $2.50 per hour from the cost and latency lesson. Wall-clock is days to a couple of weeks on a small cluster.
That is not a large bill by the standards of this course, and it is the main reason face generation was solved before general image generation: a narrow distribution and aligned data make the problem tractable at a cost a research group can afford.
The comparison with diffusion, previewed
Section 8 covers diffusion properly. The comparison to hold in mind:
| Style-based GAN | Diffusion (Section 8) | |
|---|---|---|
| Inference | 1 pass, ~20 ms | 25–50 passes, ~1 s |
| Cost per image | ~$0.000014 | ~$0.0007 |
| Training stability | Requires active management | Ordinary supervised training |
| Diversity and coverage | Weaker; mode collapse is a live risk | Stronger |
| Quality on a narrow domain | Excellent | Excellent |
| Quality on a broad domain | Poor | Excellent |
| Editability | Excellent — a structured latent space | Requires extra machinery |
Roughly a 50× inference gap. Adversarial networks did not survive out of sentiment; they survived because for narrow domains at high throughput, nothing else is close on cost.
Control and editing
A trained generator gives you random faces. A useful generator lets you get a particular face and then change one thing about it. That capability lives entirely in the structure of the latent space.
The latent space, and what interpolation proves
Take two random vectors, w₁ and w₂, generate the image for each, then generate images for a sequence of points along the straight line between them.
If the space is well-formed, you get a smooth morph: the face gradually changes identity, with every intermediate image being a plausible face. If it is not, you get plausible faces at the ends and incoherent images in between.
That smoothness is genuinely informative. It shows the generator has learned a continuous mapping from latent space to faces rather than memorising a set of outputs with noise between them. A model that memorised would have no valid intermediate points. Interpolation is therefore both a capability and a diagnostic, and it is worth running on any generator you are evaluating.
Attribute directions
The useful discovery: for many attributes, "more of this attribute" corresponds to moving along a roughly straight line in latent space.
Finding one is straightforward:
- Generate a few thousand images with their latent vectors recorded.
- Label each for the attribute — with a classifier, or by hand — for example "smiling" or "not smiling".
- Train a linear classifier on (latent vector → label).
- The direction perpendicular to that classifier's decision boundary is the smile direction.
Now add a multiple of that direction to any latent vector and the face smiles more; subtract it and it smiles less. Repeat per attribute and you have a set of sliders.
Entanglement, which is the real difficulty
The sliders are not independent, and this is the central problem of latent editing.
Move along the "smile" direction and the face may also appear younger. Move along "wears glasses" and the face may also become older and more masculine-presenting. The directions are entangled.
The cause is the data, not the model. Attributes are correlated in the training photographs: if smiling people in the dataset skew younger, then the latent direction that increases smiling also increases whatever else covaried with it. The model learned the joint distribution faithfully. Entanglement is your dataset's correlations made manipulable.
Three partial remedies:
- Use the intermediate latent space rather than the input one. As the style-based architecture above showed, the mapping network can warp the distribution, and the resulting space is measurably less entangled. This is the largest single improvement available and it is free once you have the architecture.
- Orthogonalise directions. Having found a smile direction and an age direction, project the smile direction to be perpendicular to age. Reduces the coupling; does not eliminate it, because the underlying representation is still entangled.
- Edit at specific resolution levels. Because coarse levels control pose and shape and fine levels control texture, applying an edit only at the levels where the attribute lives limits the collateral change. This is the most practical technique and it is a direct payoff from the style-based architecture.
Editing a real photograph
Users want to edit their face, not a generated one. That requires inversion: finding the latent vector whose output matches a given photograph, usually by optimising the latent to minimise a perceptual difference from the target.
The honest limitation is a trade with no clean answer. Optimise hard enough to reproduce the photograph exactly and you land outside the region of latent space where edits behave well, so the face reconstructs perfectly and edits badly. Stay in the well-behaved region and the reconstruction visibly is not quite the same person. Every product in this area is picking a point on that curve.
What makes a latent space useful
Three properties, and a space can have the third without the first two:
- Smooth — small moves cause small changes, and interpolation stays on the face manifold.
- Disentangled — a direction changes one thing.
- Invertible — a real image can be mapped back into it with acceptable fidelity.
Every generator has a latent space. Having one that satisfies these is an architectural achievement, and it is the reason the style-based design displaced its predecessors.
Safety, monitoring and follow-ups
This step is a requirement of the section, not an appendix. A system that generates photorealistic human faces is directly enabling for two serious harms, and a design that does not address them is incomplete.
The two dominant risks
Non-consensual intimate imagery and harassment. Face generation and face swapping are adjacent capabilities, and the same latent-editing machinery described above applies to a real person's inverted photograph. This is the harm with the most documented victims, it falls overwhelmingly on women, and it is the reason identity controls below are not optional.
Impersonation and synthetic identities. Generated faces have been used at scale for fake social media profiles, fraudulent professional accounts, and coordinated influence operations — a well-documented and repeatedly-reported pattern. The generated face is valuable to an attacker precisely because a reverse image search returns nothing.
The controls, as an architecture
Design these as components with costs, exactly as in Safety as architecture:
1. Identity resemblance checking. Before returning an image, embed the generated face with a face-recognition model and compare against an index of public figures. Above a similarity threshold, discard and regenerate. Adds roughly 10 ms and a small index. It catches accidental resemblance to well-known people; it cannot cover the general population, and saying so is the honest position.
2. Refuse the identity-targeted path entirely. Do not accept a reference photograph in this product. The moment a user can supply a target face, you are building Section 10's product and you need its consent verification.
3. Invisible watermarking and provenance credentials. Embed a signal in the pixels, and attach signed provenance metadata. Both are good-faith labelling: watermarks survive re-compression and mild cropping reasonably well and are removable by a motivated adversary; metadata is stripped by a screenshot. Do both, and do not claim either is a defence.
4. Access control and rate limiting. Verified accounts for API access, per-account rate limits, and retained generation logs with a traceable identifier. Most abuse at scale needs volume, and volume is the thing you can actually see.
5. Human review on a sample, with a report path and a runbook.
Detection, and honesty about it
Classifiers that detect generated faces exist and work reasonably on the generator they were trained against. They generalise poorly to new generators and degrade under compression, resizing, and re-encoding — the exact things that happen to any image posted online.
Simple heuristics have also worked and then stopped working. The aligned-eye position from the data lesson is the clearest example: it was a reliable tell for one generation of aligned-face models and is easily defeated by re-cropping.
The correct claim: detection is a useful signal in a layered system and is not a reliable gate. A candidate who says detection solves deepfakes has said something false; a candidate who says provenance-by-construction (watermarking plus credentials at generation time) is more durable than detection-after-the-fact is saying the right thing.
Monitoring
Track output diversity continuously — a drift towards a narrower distribution is mode collapse appearing in production. Track the resemblance-check rejection rate, since a rise means either drift or an attack. Track per-group FID against the balanced reference set from the data lesson on every model update. And sample generated output for human review, including specifically for the defect classes in the metrics lesson.
The legitimate uses that justify the technology
Worth stating, because a good answer is not only a list of harms:
- Privacy-preserving synthetic data. Replace real faces in a dataset with generated ones, so a detection or blurring model can be trained without holding photographs of real people. This is a direct privacy improvement and it is the strongest justification for the technology.
- De-identification. Replace faces in published imagery — street-level mapping, medical teaching material — with plausible synthetic ones instead of blurring, preserving downstream usefulness.
- Avatars and creative tools, where users represent themselves without using their own face.
- Data augmentation for under-represented groups, which is genuinely useful and carries the obvious circularity: a generator trained on skewed data cannot manufacture the diversity it never saw.