Deep Learning Essentials

Course Content

Deep Learning Essentials

13 sections · 61 lessons

What are GANs?


One GAN training stepRandomnoise,16 numbersGenerator makesa fake sampleDiscriminatorscores realvs fakeD learns: realis 1, fake is 0G learns fromD's gradientto fool itfake.detach() keeps D's update from changing G.
Only the generator is kept at the end; the discriminator exists to turn looks-fake into a gradient the generator can follow.

What you need to know

The two players

  • Generator G — input: a random noise vector, say 16 numbers. Output: a fake sample, such as an image or a table row.
  • Discriminator D — input: a sample. Output: the probability it is real.

One training step

Python
import torch, torch.nn as nnbce = nn.BCEWithLogitsLoss()G = nn.Sequential(nn.Linear(16, 64), nn.ReLU(), nn.Linear(64, 30))          # noise -> fake rowD = nn.Sequential(nn.Linear(30, 64), nn.LeakyReLU(0.2), nn.Linear(64, 1))   # row -> real/fake logitopt_g = torch.optim.Adam(G.parameters(), lr=2e-4, betas=(0.5, 0.999))opt_d = torch.optim.Adam(D.parameters(), lr=2e-4, betas=(0.5, 0.999))def train_step(real):    n = real.size(0)    fake = G(torch.randn(n, 16))    # 1. Discriminator: label real as 1, fake as 0    d_loss = bce(D(real), torch.ones(n, 1)) + bce(D(fake.detach()), torch.zeros(n, 1))    opt_d.zero_grad(); d_loss.backward(); opt_d.step()    # 2. Generator: try to make D say 1 on the fakes    g_loss = bce(D(fake), torch.ones(n, 1))    opt_g.zero_grad(); g_loss.backward(); opt_g.step()

fake.detach() stops the discriminator's loss from updating the generator. In step 2, the gradient flows back through D into G, telling G how to change its output so D is more convinced. This is the minimax game: D maximises its accuracy, G minimises it.

Known problems

  • Training instability — the two losses oscillate; if D becomes too strong, G's gradients vanish.
  • Mode collapse — G finds a few outputs that fool D and produces only those, losing variety. A face GAN that makes the same five faces has collapsed.
  • Hard to evaluate — no simple likelihood. People use metrics like FID, which compare feature statistics of real and generated images.

Fixes such as Wasserstein loss, spectral normalisation and progressive growing made GANs like StyleGAN produce very realistic faces.

Where GANs stand today

Diffusion models now lead image, video and audio generation because they train stably and cover more variety. GANs still matter where speed is critical, because a GAN makes an image in one forward pass while a diffusion model needs many steps. Adversarial losses are also used inside other systems, such as audio vocoders and to speed up diffusion models.

A real-life example

A bank's fraud dataset has 500,000 normal transactions and only 800 fraud cases. The fraud classifier struggles because it sees so few fraud examples.

The team trains a tabular GAN (CTGAN is a common library for this) on the 800 fraud rows only, then generates 5,000 synthetic fraud rows. They check quality by training a classifier to separate real from synthetic fraud: if it can barely do better than chance, the synthetic rows are realistic.

Adding the synthetic rows to training raises validation fraud recall from 0.71 to 0.78, measured only on real transactions. They watch for mode collapse: the first GAN produced almost only ₹49,999 purchases at electronics stores, which is one fraud pattern repeated. Lowering the learning rate and training longer produced more variety.

Follow-up questions to expect

  • "What is mode collapse?" — The generator produces only a few kinds of output that fool the discriminator, instead of covering the full variety of real data.
  • "GAN vs VAE?" — GANs give sharper samples but train unstably; VAEs train stably with a clear loss but samples tend to be blurrier.
  • "Why have diffusion models replaced GANs for images?" — They train with a simple, stable denoising loss and cover the data distribution better, at the cost of slower sampling.