Course Content
Deep Learning Essentials
13 sections · 61 lessons
What is the difference between dropout and batch normalization?
What you need to know
Dropout, with numbers
With drop rate p = 0.5, each activation is kept or zeroed at random, and the survivors are scaled up by 1 / (1 − p) = 2 so the expected total stays the same. This is called inverted dropout, and it is what PyTorch does.
activations [2, 4, 6, 8]random mask [1, 0, 1, 0]output [4, 0, 12, 0] # survivors × 2A different mask is drawn every step. No neuron can count on a particular partner being present, so each has to learn features that are useful on their own. At inference, dropout does nothing: all neurons are used and no scaling is needed, because it was done during training.
Batch normalisation, with numbers
For one feature across a batch of 4 examples:
values [2, 4, 6, 8]batch mean 5batch variance 5 (std ≈ 2.236)normalised [−1.34, −0.45, 0.45, 1.34]output γ × normalised + β # γ, β are learnedEvery layer's inputs now stay in a stable range, whatever the earlier layers are doing. That lets you use larger learning rates and makes training less sensitive to initialisation. Because each example's output depends slightly on which other examples share its batch, BN also adds a little noise, which regularises mildly.
At inference there may be a single example, so a batch mean makes no sense. BN instead uses a running mean and variance collected during training.
Dropout
- Goal: reduce overfitting
- Randomly zeroes activations, scales survivors
- No learnable parameters, one hyperparameter p
- Inference: turned off completely
Batch normalisation
- Goal: stable, faster training
- Normalises with batch mean and variance
- Learns scale γ and shift β per feature
- Inference: uses stored running statistics
Both depend on train and eval mode
model.train() turns dropout on and makes BN use batch statistics. model.eval() turns dropout off and makes BN use running statistics. Forgetting model.eval() before validation is one of the most common bugs in PyTorch code.
Using them together
Putting dropout right before a BN layer can cause a mismatch: BN's running variance is computed with dropout's noise during training, but at inference the noise is gone. The usual convention in CNNs is conv, BN, ReLU in each block, with dropout only in the classifier head, or no dropout at all. Transformers use layer normalisation instead of BN, normalising across features within each example, plus dropout.
A real-life example
A medical team trains a 3D CNN on CT scans. Each scan is large, so only 2 fit on the GPU at once. With batch normalisation, the "batch mean" is the average of just 2 scans, so it jumps around wildly between steps, and validation accuracy swings by 10 points between epochs.
They replace BatchNorm3d with GroupNorm, which normalises within each example over groups of channels, so it does not depend on batch size. Training becomes smooth. They keep dropout of 0.3 in the classifier head to reduce overfitting on their 900 scans. The two layers were never alternatives: one fixed stability, the other fought overfitting.
Follow-up questions to expect
- "Why does batch norm fail with small batches?" — The mean and variance of 2 or 4 examples are poor estimates, so normalisation adds a lot of noise, and the running statistics used at inference do not match training. Group or layer norm avoids this.
- "Where do you put batch norm, before or after the activation?" — The original paper put it before the activation, and conv, BN, ReLU is the common convention. Both orders appear in practice.
- "Why is dropout less common in modern CNNs?" — BN, augmentation and weight decay already regularise, and dropout in conv layers interacts badly with BN. It survives in classifier heads and in transformers.