Course Content
LLMs Deep Dive
10 sections · 40 lessons
What is the vanishing gradient problem, and how do Transformers solve it?
What you need to know
Why the signal shrinks
The gradient at an early layer is a product of one factor per layer (or per time step). If each factor is below 1, the product collapses:
Sigmoid's maximum slope = 0.2510 sigmoid layers: 0.25^10 ≈ 0.000001RNN, factor 0.9 per step, 50 steps: 0.9^50 ≈ 0.005The early layers get a millionth of the signal of the late ones, so they stay near their random start. If factors are above 1, the same maths gives exploding gradients.
What transformers do differently
- Direct paths across positions — attention links token 1 and token 500 in one step, so the gradient does not travel through 499 recurrent updates.
- Residual connections — each sub-layer computes x + F(x). Its derivative is 1 + F'(x), so even if F' is tiny, the gradient passes through with a factor of about 1.
- Normalisation — LayerNorm or RMSNorm rescales activations; placing it before each sub-layer (pre-norm) is what makes very deep stacks stable.
- Non-saturating activations — GELU and SwiGLU keep useful slopes for positive inputs, unlike sigmoid and tanh which flatten at both ends.
- Training hygiene — scaled initialisation, learning-rate warmup and gradient clipping.
What transformers do not solve
- Exploding gradients and loss spikes still happen in large-scale training; clipping, lower learning rates and restarting from a checkpoint are standard.
- Long-context learning is limited by data and position handling, not by vanishing gradients.
LSTMs as a partial fix
LSTMs added gates and a cell state with an additive update, which carried gradients further than plain RNNs — an early form of the same "additive path" idea residuals use.
A real-life example
A team trains a small transformer from scratch for Hindi–English code-mixed intent detection, with 24 layers. They copy an old configuration that places LayerNorm after the residual add (post-norm) and use no warmup. Training loss stays flat at the starting value for 5,000 steps. Per-layer gradient norms show the bottom 8 layers receiving about a thousandth of the gradient of the top layers.
Switching to pre-norm and adding 1,000 warmup steps gets the loss falling within the first few hundred steps. The model itself did not change; the path the gradient could travel did.
Follow-up questions to expect
- "How do you detect vanishing gradients?" — Log gradient norms per layer; early layers near zero while late layers are normal is the signature.
- "What is pre-norm versus post-norm?" — Pre-norm normalises the input to each sub-layer and leaves the residual path clean; post-norm normalises after the add. Pre-norm is easier to train deep; almost all modern LLMs use it.
- "Why did ReLU help?" — Its slope is exactly 1 for positive inputs, so it does not shrink gradients the way sigmoid does.