Course Content
Deep Learning Essentials
13 sections · 61 lessons
What is dropout, and how does it reduce overfitting?
What you need to know
What happens to the numbers
With dropout probability p = 0.5, each activation is kept with probability 0.5. Kept values are multiplied by 1 / (1 − p) = 2. This is called inverted dropout, and it is what PyTorch does.
1import torch, torch.nn as nn23drop = nn.Dropout(p=0.5)4x = torch.ones(8)56drop.train(); print(drop(x)) # e.g. tensor([0., 0., 2., 0., 0., 0., 2., 2.])7drop.eval(); print(drop(x)) # tensor([1., 1., 1., 1., 1., 1., 1., 1.])In training mode, each value is zeroed with probability 0.5 (here five of the eight) and the survivors become 2, so the average stays near 1. In eval mode nothing changes. Because the scaling happens during training, no correction is needed at inference.
A new random mask is drawn for every example in every forward pass, so each step trains a different "thinned" network.
Why it reduces overfitting
- Breaks co-adaptation. Without dropout, neuron A may learn to fix neuron B's specific mistakes. Those fragile partnerships often encode noise. With dropout, B may be absent at any step, so A must learn something useful by itself.
- Implicit ensemble. A layer of
nneurons has2^npossible thinned versions. Training with dropout trains a huge family of them with shared weights. Using the full network at test time roughly averages their predictions, and averaging reduces variance. - Noise as regularisation. The model cannot fit any single training example exactly, because its own computation changes every step.
Practical settings
- Fully connected layers:
p = 0.2to0.5. - Convolutional layers: usually small or none; batch norm, augmentation and weight decay do more there.
- Transformers: often
p = 0.1on attention and feed-forward outputs. Many large LLM pretraining runs use no dropout at all, because they have so much data. - Dropout usually makes training loss higher and training slower. That is expected.
A real-life example
A bank trains a fraud classifier on 50,000 card transactions, of which 600 are fraud. Their 3-layer MLP reaches 0.97 recall on training fraud cases but 0.68 on validation. Some neurons have learned combinations like "merchant ID 88213 and amount ₹4,999", which match exactly a few training frauds and nothing else.
Adding nn.Dropout(0.3) after each hidden layer means those very specific combinations are often broken during training. The model is pushed towards broader signals such as "new device and unusual hour and high amount". Training recall falls to 0.90, validation recall rises to 0.81. A smaller gap, and a better model.
The team's first deployment had a bug: the serving code never called model.eval(), so dropout stayed on in production. The same transaction got different fraud scores on each call. Calling model.eval() fixed it.
Follow-up questions to expect
- "Why scale by 1/(1−p)?" — So the expected activation during training matches inference. Otherwise the next layer would see inputs twice as large at test time.
- "Can dropout be used at inference?" — Yes, on purpose: Monte Carlo dropout runs many forward passes with dropout on and uses the spread of predictions as an uncertainty estimate.
- "Dropout with batch norm?" — They can conflict, because dropout changes activation variance between training and inference, which upsets batch norm's statistics. If both are used, dropout is usually placed after the last batch-norm layer.