Course Content
Deep Learning Essentials
13 sections · 61 lessons
What is the vanishing gradient problem, and how is it different from the exploding gradient problem?
What you need to know
Why a product of many numbers is fragile
The gradient for a weight in layer 1 of a 10-layer network is a product of about ten factors, one per layer. Each factor combines the activation's derivative and the layer's weights.
- Multiply ten numbers around 0.25 and you get about 0.00000095. The first layer's gradient is a millionth of the last layer's.
- Multiply fifty numbers around 1.5 and you get about 640 million.
- Even 0.9 repeated fifty times gives 0.005.
Nothing is broken in the maths. Repeated multiplication is simply unstable unless each factor stays close to 1.
Why sigmoid makes vanishing worse
The sigmoid's derivative is at most 0.25, at input 0, and near zero for inputs beyond about ±5. So every sigmoid layer multiplies the gradient by 0.25 at best. Tanh peaks at 1 but also flattens at the extremes. This is why deep sigmoid networks were so hard to train before ReLU.
Why RNNs get both problems
An RNN applies the same weight matrix at every time step. Backpropagation through 100 time steps multiplies by that matrix 100 times. If its largest scaling factor is a bit below 1, the gradient from 100 steps ago vanishes and the network cannot learn long-range patterns. A bit above 1, it explodes. That is why LSTMs and GRUs were invented, and why gradient clipping is standard for RNNs.
How to tell which one you have
Vanishing
- Loss drops a little, then goes flat
- Early layers' gradient norms near zero
- Early layers' weights barely change
- Model ignores long-range context
Exploding
- Loss spikes or becomes
NaN - Gradient norms suddenly huge
- Weights jump to very large values
- Often triggered by one unusual batch
You can check directly by printing gradient norms per layer after loss.backward():
for name, p in model.named_parameters(): if p.grad is not None: print(f"{name:30s} {p.grad.norm().item():.2e}")If the first layers show 1e-07 while the last show 1e-01, gradients are vanishing.
A real-life example
A retail chain forecasts daily sales for each store with a plain RNN over the last 365 days. Sales jump every year around Diwali, so the right signal for predicting this October is last October, about 365 steps back.
The RNN learns the weekly pattern (7 steps back) well but completely misses the festival spike. Printing gradient norms shows why: the gradient reaching step 300 is smaller than 1e-10. The error signal from last October never arrives. Switching to an LSTM, whose cell state carries information forward with additive updates, and adding a "days to Diwali" feature fixes the forecast. On a different run with a higher learning rate, the same RNN hits NaN after one batch containing a store's clearance-sale day with 40 times normal sales, an exploding-gradient case that clipping would have caught.
Follow-up questions to expect
- "Why do gradients vanish in early layers and not late ones?" — Late layers are close to the loss, so their gradient passes through few multiplications. Early layers are at the end of a long chain of factors.
- "Can a network have both problems?" — Yes. An RNN can vanish over long spans and explode on certain inputs, and a badly initialised deep network can vanish in some layers and explode in others.
- "Is vanishing gradient the same as a dead ReLU?" — Related but different. A dead ReLU gives an exact zero gradient for one unit; vanishing gradients are a gradual shrinking across many layers.