Deep Learning Essentials

Course Content

Deep Learning Essentials

13 sections · 61 lessons

What is the vanishing gradient problem, and how is it different from the exploding gradient problem?


Gradient left after each sigmoid layer, best case10.250.06250.01560.003901234layernext to lossfour layers backEach layer multiplies by at most 0.25; ten layers leaves about one-millionth.
Early layers learn slowest because their gradient is the product of every factor between them and the loss.

What you need to know

Why a product of many numbers is fragile

The gradient for a weight in layer 1 of a 10-layer network is a product of about ten factors, one per layer. Each factor combines the activation's derivative and the layer's weights.

  • Multiply ten numbers around 0.25 and you get about 0.00000095. The first layer's gradient is a millionth of the last layer's.
  • Multiply fifty numbers around 1.5 and you get about 640 million.
  • Even 0.9 repeated fifty times gives 0.005.

Nothing is broken in the maths. Repeated multiplication is simply unstable unless each factor stays close to 1.

Why sigmoid makes vanishing worse

The sigmoid's derivative is at most 0.25, at input 0, and near zero for inputs beyond about ±5. So every sigmoid layer multiplies the gradient by 0.25 at best. Tanh peaks at 1 but also flattens at the extremes. This is why deep sigmoid networks were so hard to train before ReLU.

Why RNNs get both problems

An RNN applies the same weight matrix at every time step. Backpropagation through 100 time steps multiplies by that matrix 100 times. If its largest scaling factor is a bit below 1, the gradient from 100 steps ago vanishes and the network cannot learn long-range patterns. A bit above 1, it explodes. That is why LSTMs and GRUs were invented, and why gradient clipping is standard for RNNs.

How to tell which one you have

Vanishing

  • Loss drops a little, then goes flat
  • Early layers' gradient norms near zero
  • Early layers' weights barely change
  • Model ignores long-range context

Exploding

  • Loss spikes or becomes NaN
  • Gradient norms suddenly huge
  • Weights jump to very large values
  • Often triggered by one unusual batch

You can check directly by printing gradient norms per layer after loss.backward():

Python
for name, p in model.named_parameters():    if p.grad is not None:        print(f"{name:30s} {p.grad.norm().item():.2e}")

If the first layers show 1e-07 while the last show 1e-01, gradients are vanishing.

A real-life example

A retail chain forecasts daily sales for each store with a plain RNN over the last 365 days. Sales jump every year around Diwali, so the right signal for predicting this October is last October, about 365 steps back.

The RNN learns the weekly pattern (7 steps back) well but completely misses the festival spike. Printing gradient norms shows why: the gradient reaching step 300 is smaller than 1e-10. The error signal from last October never arrives. Switching to an LSTM, whose cell state carries information forward with additive updates, and adding a "days to Diwali" feature fixes the forecast. On a different run with a higher learning rate, the same RNN hits NaN after one batch containing a store's clearance-sale day with 40 times normal sales, an exploding-gradient case that clipping would have caught.

Follow-up questions to expect

  • "Why do gradients vanish in early layers and not late ones?" — Late layers are close to the loss, so their gradient passes through few multiplications. Early layers are at the end of a long chain of factors.
  • "Can a network have both problems?" — Yes. An RNN can vanish over long spans and explode on certain inputs, and a badly initialised deep network can vanish in some layers and explode in others.
  • "Is vanishing gradient the same as a dead ReLU?" — Related but different. A dead ReLU gives an exact zero gradient for one unit; vanishing gradients are a gradual shrinking across many layers.