Deep Learning Essentials

Course Content

Deep Learning Essentials

13 sections · 61 lessons

What is dropout, and how does it reduce overfitting?


Eight activations of 1.0 through Dropout(p=0.5), training mode0020002201234567droppedkept,scaled by 2In eval mode the same eight values pass through as 1.0: no mask, no scaling.
Scaling survivors by 1/(1 − p) during training keeps the expected value the same, which is why nothing needs correcting at inference.

What you need to know

What happens to the numbers

With dropout probability p = 0.5, each activation is kept with probability 0.5. Kept values are multiplied by 1 / (1 − p) = 2. This is called inverted dropout, and it is what PyTorch does.

Python
import torch, torch.nn as nndrop = nn.Dropout(p=0.5)x = torch.ones(8)drop.train(); print(drop(x))   # e.g. tensor([0., 0., 2., 0., 0., 0., 2., 2.])drop.eval();  print(drop(x))   # tensor([1., 1., 1., 1., 1., 1., 1., 1.])

In training mode, each value is zeroed with probability 0.5 (here five of the eight) and the survivors become 2, so the average stays near 1. In eval mode nothing changes. Because the scaling happens during training, no correction is needed at inference.

A new random mask is drawn for every example in every forward pass, so each step trains a different "thinned" network.

Why it reduces overfitting

  • Breaks co-adaptation. Without dropout, neuron A may learn to fix neuron B's specific mistakes. Those fragile partnerships often encode noise. With dropout, B may be absent at any step, so A must learn something useful by itself.
  • Implicit ensemble. A layer of n neurons has 2^n possible thinned versions. Training with dropout trains a huge family of them with shared weights. Using the full network at test time roughly averages their predictions, and averaging reduces variance.
  • Noise as regularisation. The model cannot fit any single training example exactly, because its own computation changes every step.

Practical settings

  • Fully connected layers: p = 0.2 to 0.5.
  • Convolutional layers: usually small or none; batch norm, augmentation and weight decay do more there.
  • Transformers: often p = 0.1 on attention and feed-forward outputs. Many large LLM pretraining runs use no dropout at all, because they have so much data.
  • Dropout usually makes training loss higher and training slower. That is expected.

A real-life example

A bank trains a fraud classifier on 50,000 card transactions, of which 600 are fraud. Their 3-layer MLP reaches 0.97 recall on training fraud cases but 0.68 on validation. Some neurons have learned combinations like "merchant ID 88213 and amount ₹4,999", which match exactly a few training frauds and nothing else.

Adding nn.Dropout(0.3) after each hidden layer means those very specific combinations are often broken during training. The model is pushed towards broader signals such as "new device and unusual hour and high amount". Training recall falls to 0.90, validation recall rises to 0.81. A smaller gap, and a better model.

The team's first deployment had a bug: the serving code never called model.eval(), so dropout stayed on in production. The same transaction got different fraud scores on each call. Calling model.eval() fixed it.

Follow-up questions to expect

  • "Why scale by 1/(1−p)?" — So the expected activation during training matches inference. Otherwise the next layer would see inputs twice as large at test time.
  • "Can dropout be used at inference?" — Yes, on purpose: Monte Carlo dropout runs many forward passes with dropout on and uses the spread of predictions as an uncertainty estimate.
  • "Dropout with batch norm?" — They can conflict, because dropout changes activation variance between training and inference, which upsets batch norm's statistics. If both are used, dropout is usually placed after the last batch-norm layer.