Course Content
Introduction to Generative AI
3 sections · 9 lessons
Overview of Generative Architectures
Here is the obvious way to build a generative model. You want p(x), the probability of an image. So build a network that takes an image and outputs a number, train it to output large numbers for real images and small numbers for everything else, and then sample from it.
Try to actually do this and you hit a wall within about a minute. For those outputs to be probabilities they must sum to one across every possible image. So you need
and Z — the normalising constant — requires evaluating your network on every image that could exist. For a modest 256×256 colour image that is 256196608 evaluations. Not slow. Not expensive. Physically impossible, by a margin so large that no amount of hardware will ever matter.
And even if you somehow had p(x), you would still be stuck, because knowing a probability for each item does not tell you how to draw one from an unimaginably large space.
Every generative architecture in use is a different escape route from this single problem. That is the most useful way to hold them in your head — not as four unrelated inventions, but as four answers to one impossible integral.
| Family | The dodge |
|---|---|
| Autoregressive (transformers) | Split the hard distribution into many easy ones, each normalised over a small set |
| Variational autoencoders | Give up on the exact likelihood; optimise a provable lower bound instead |
| Generative adversarial networks | Abandon likelihood entirely; train a second network to judge realism |
| Diffusion models | Never make the hard jump; take thousands of tiny, easy steps |
You cannot normalise a distribution over all possible images. Every architecture here is a way of not having to.
Autoregressive models: make the problem small
The chain rule of probability is exact and free:
Each factor is a distribution over one element, not over the whole object. If the element is a word drawn from a vocabulary of 50,000, then the normalising sum has 50,000 terms — a rounding error of compute. The impossible integral has been replaced by a long series of trivial ones.
A transformer is the machinery that makes each conditional accurate. Its core operation, self-attention, lets every position look at every earlier position and weight them by relevance. Predicting the last word of "The trophy would not fit in the suitcase because it was too large" requires knowing that it refers to the trophy — a dependency 11 words back. Attention makes that a direct lookup rather than something that has to survive being passed hand to hand through a chain.
Generation is then a loop: predict a distribution over the next element, sample one, append it, repeat.
1tokens = tokenise(prompt)2for _ in range(max_new_tokens):3 logits = model(tokens)[-1] # scores for every vocabulary entry4 probs = softmax(logits / temperature)5 tokens.append(sample(probs)) # one element per forward passWhat you get. An exact likelihood — you can ask the model precisely how probable any given sequence is, which almost nothing else here can do. Training parallelises beautifully: every position's prediction is computed at once, since the correct earlier elements are already known.
What it costs. Generation is strictly sequential. One thousand tokens means roughly one thousand forward passes, and no hardware fixes that, because token 500 cannot be computed before token 499 exists. Serving tricks such as speculative decoding let a small, fast model guess several tokens ahead and the large model check them all in one pass, which helps, but the dependency on earlier tokens remains. There is also exposure bias: during training the model always conditions on correct history, but at generation time it conditions on its own output, so an early mistake becomes a permanent premise it will loyally build on.
Variational autoencoders: settle for a bound
Start with a compression idea. An encoder squeezes an image down to a few dozen numbers — the latent code — and a decoder expands it back. Train the pair so the output matches the input. That is a plain autoencoder, and it compresses well.
It also cannot generate. Feed the decoder a random code and you get garbage, because the encoder was free to scatter its codes anywhere it liked. Nothing forced the occupied regions to join up. The space between two valid codes is mostly empty, and empty means noise.
The variational fix is to make the encoder output a small cloud rather than a point — a mean μ and a spread σ — and add a penalty pulling those clouds towards a standard bell curve centred at the origin. Now the clouds overlap and cover the region, so a random draw from that bell curve lands somewhere the decoder understands.
This is derived, not invented. The true objective logp(x) is intractable for the reason above, but it can be bounded from below:
That right-hand side is the evidence lower bound, or ELBO. Maximising it pushes up the thing you actually wanted. The two terms pull against each other: reconstruction wants every input to get its own distinct, precise code; the KL term wants all codes to look like the same anonymous bell curve. The balance point is a latent space that is both informative and continuous.
What you get. Fast single-pass generation, stable training that essentially always converges, and a genuinely useful latent space — interpolate between two codes and you get a smooth morph between two outputs, which is why VAEs remain the tool of choice for representation learning.
What it costs. Blur. Because the model averages over an entire cloud of codes and typically optimises squared error, it hedges — and the mathematically optimal hedge between several sharp possibilities is a smeared average of them. VAE samples look like photographs seen through frosted glass.
Generative adversarial networks: let a critic decide
Blur comes from averaging, and averaging comes from likelihood objectives. So drop the likelihood.
A GAN runs two networks against each other. The generator maps random noise to an image. The discriminator looks at an image and judges whether it came from the real dataset or from the generator. They train simultaneously with opposite goals:
The discriminator's job is a plain classification problem, which networks are excellent at. The generator's gradient comes entirely from the discriminator's verdict, so it is not told "match this target" — it is told "you were caught, here is the direction that would have fooled me". No averaging over possibilities is ever performed, so nothing forces the output towards a blurry mean. The result was, for years, unmatched sharpness.
What it costs. Two things, and both are serious.
Instability. There is no single loss decreasing towards a minimum. There is an equilibrium between two moving opponents, and it can oscillate or collapse. If the discriminator gets too good too fast, it rejects everything with total confidence, the generator's gradient goes to zero, and training dies.
Mode collapse. Nothing in the objective requires covering the data. A generator that produces one excellent image of a golden retriever, over and over, fools the discriminator perfectly on that image. It has won the game and lost the task. Diversity was never in the loss function.
A GAN is scored on whether each sample looks real, never on whether the samples together look like the dataset. Every characteristic GAN failure follows from that omission.
Diffusion models: many easy steps
The last escape route is the one that ended up winning for images, and its idea is almost embarrassingly simple: the jump from noise to a photograph is impossible in one move, so do not make it in one move.
Forward process. Take a real image and add a small amount of Gaussian noise. Repeat, perhaps a thousand times. The image dissolves into pure static. This direction requires no learning at all — it is a fixed recipe, and there is a closed form that jumps straight to any step, so training does not have to simulate the whole chain.
Reverse process. Train a network to look at a noisy image and predict which noise was added. Then generate by starting from pure static and applying that prediction repeatedly, stripping a little noise each time until an image emerges.
The reason this works is that each individual step is easy. Removing 0.1% of the noise from an image that is 40% noise is a mild, well-conditioned regression problem — nothing like the intractable leap from static to a photograph. And unlike a GAN, the training signal is a plain regression loss against a known target, so there is no adversarial equilibrium to lose.
1# Training: one noise level per example, sampled at random2t = randint(1, T) # a random step3noise = randn_like(x0)4x_t = sqrt(alpha_bar[t]) * x0 + sqrt(1 - alpha_bar[t]) * noise5loss = mse(model(x_t, t), noise) # predict the noise that was added67# Sampling: start from static, walk back down8x = randn(shape)9for t in reversed(range(T)):10 x = denoise_step(x, model(x, t), t)What you get. The best image quality available, stable training, and full mode coverage — because the objective is a likelihood-style bound, the model is penalised for ignoring parts of the data, so it cannot collapse onto one output the way a GAN can.
What it costs. Speed. One sample means dozens to hundreds of network evaluations where a GAN or VAE needs one. This is why so much recent work is about better samplers and distillation — squeezing 1,000 steps down to 20, or to 4.
A close relative: flow matching
Many of the newest image and video models, including Stable Diffusion 3 and FLUX.1, are trained with flow matching (one popular form is called rectified flow). The set-up is the same: blend a real image with noise, and train a network on the blend. The difference is the target. Instead of predicting the noise, the network predicts a velocity: which direction to move, and how fast, along a straight line between noise and image. Sampling follows that line from noise back to an image, and straighter paths can be followed accurately in fewer steps. For this course, treat flow models as members of the diffusion family: same strengths, same speed problem, somewhat smaller.
Diffusion has even reached text. Research systems such as Google's experimental Gemini Diffusion, and commercial ones such as Inception's Mercury, generate a whole block of text by refining it over several passes rather than one token at a time. They are fast, but the leading general-purpose language models are still mostly autoregressive.
The trilemma
Lay the four families against the three things anyone wants and a pattern appears.
| Sample quality | Coverage of the data | Sampling speed | Training stability | |
|---|---|---|---|---|
| Autoregressive | Excellent | Excellent | Poor — one pass per element | Excellent |
| VAE | Poor — blurry | Good | Excellent — one pass | Excellent |
| GAN | Excellent | Poor — mode collapse | Excellent — one pass | Poor |
| Diffusion | Excellent | Excellent | Poor — many passes | Good |
Nothing has all three of quality, coverage and speed. This is sometimes called the generative learning trilemma, and it is not an accident of engineering — it reflects the fact that each family bought its way out of the normalising constant with a different currency. Autoregressive models paid in sequential time. VAEs paid in sharpness. GANs paid in stability and coverage. Diffusion models paid in compute per sample.
| Also worth knowing | AR | VAE | GAN | Diffusion |
|---|---|---|---|---|
| Exact likelihood available | Yes | Lower bound only | No | Lower bound only |
| Useful latent space | No | Yes | Partly | Not natively |
| Can score an unseen input | Yes | Approximately | No | Approximately |
| Easy to condition on text | Yes | Yes | Harder | Yes |
| Best natural fit | Discrete sequences | Compression, representations | Fast image synthesis | High-fidelity continuous data |
The hybrids that actually ship
Almost no production system is one pure family. The interesting engineering is in combinations that let one architecture cover another's weakness.
Latent diffusion: VAE plus diffusion
Running diffusion directly on 512×512×3 pixels is ruinously expensive — every one of hundreds of steps operates on 786,432 numbers. So train a VAE to compress images into a small latent grid, roughly 64×64×4, run the entire diffusion process there, and decode once at the end.
The arithmetic is decisive: about 48 times fewer numbers per step. This one change is why text-to-image generation moved from research clusters to consumer graphics cards. It is also a clean division of labour — the VAE handles the high-frequency detail it is good at, and diffusion handles the semantic structure it is good at, so the VAE's characteristic blur never appears.
Discrete tokens plus a transformer
A vector-quantised VAE compresses an image or an audio clip into a grid of discrete codes drawn from a learned codebook. Once the data is a sequence of discrete symbols, a transformer can model it with exactly the machinery used for text. This is the standard route for audio and music generation, and it means one architecture serves every modality once the encoder has done its job.
Diffusion transformers
Early diffusion models used a convolutional U-Net as the denoiser. Replacing it with a transformer that treats latent patches as tokens gives up some built-in image structure and gains something more valuable: transformers scale predictably with parameters and data, in a way convolutional stacks do not. Current large image and video systems are largely built this way.
Adversarial distillation
Train a diffusion model normally, then train a fast student to match its output in a handful of steps, using a discriminator to keep the student's samples sharp. Diffusion supplies coverage and quality; the adversarial component supplies speed. The trilemma is not solved, but two of its corners can be borrowed at once.
What this means when you build something
Choosing an architecture is really choosing which corner of the trilemma you can afford to lose.
| Your situation | Family to reach for | Because |
|---|---|---|
| Data is discrete sequences — text, code, symbols | Autoregressive transformer | Chain-rule factorisation is exact and natural here |
| You need the highest possible image or video fidelity | Diffusion, in a latent space | Best quality with full coverage; cost is per-sample compute |
| Real-time generation on modest hardware | GAN, or a distilled diffusion student | Single forward pass |
| You want a latent space to search, interpolate or cluster in | VAE | Only family that gives a well-behaved encoder for free |
| You must detect anomalies or score inputs | Autoregressive, or a VAE | They can report how likely an input is; GANs cannot |
| Training budget is small and failure is unacceptable | VAE or diffusion | GAN training can fail outright and give you nothing |
Two warnings from practice. First, GAN training is not merely harder — it can fail in a way that produces no usable model at all after a full training run, which is a project risk rather than a tuning inconvenience. If you cannot afford that outcome, pick something with a monotone loss curve.
Second, resist the pull towards a single "best" architecture. The strongest systems are compositions, and the reason is visible in the trilemma table: each family's weakness is another family's strength. Latent diffusion exists because someone noticed a VAE's compression fixed a diffusion model's cost problem while the diffusion model fixed the VAE's blur. That is the pattern to look for — not which architecture wins, but which two cancel each other's failure mode.
Check your understanding
0 of 3 answered
1.Why can't you simply train a network to output p(x) for any image and sample from it?
2.A team needs thousands of product images a day on a small GPU budget and cannot risk a training run that yields nothing. Which choice fits best?
3.Why is the latent diffusion hybrid more than the sum of its parts?