LLMs Deep Dive

Course Content

LLMs Deep Dive

10 sections · 40 lessons

What is the vanishing gradient problem, and how do Transformers solve it?


What you need to know

Why the signal shrinks

The gradient at an early layer is a product of one factor per layer (or per time step). If each factor is below 1, the product collapses:

Text
Sigmoid's maximum slope = 0.2510 sigmoid layers:  0.25^10 ≈ 0.000001RNN, factor 0.9 per step, 50 steps:  0.9^50 ≈ 0.005

The early layers get a millionth of the signal of the late ones, so they stay near their random start. If factors are above 1, the same maths gives exploding gradients.

What transformers do differently

  1. Direct paths across positions — attention links token 1 and token 500 in one step, so the gradient does not travel through 499 recurrent updates.
  2. Residual connections — each sub-layer computes x + F(x). Its derivative is 1 + F'(x), so even if F' is tiny, the gradient passes through with a factor of about 1.
  3. Normalisation — LayerNorm or RMSNorm rescales activations; placing it before each sub-layer (pre-norm) is what makes very deep stacks stable.
  4. Non-saturating activations — GELU and SwiGLU keep useful slopes for positive inputs, unlike sigmoid and tanh which flatten at both ends.
  5. Training hygiene — scaled initialisation, learning-rate warmup and gradient clipping.

What transformers do not solve

  • Exploding gradients and loss spikes still happen in large-scale training; clipping, lower learning rates and restarting from a checkpoint are standard.
  • Long-context learning is limited by data and position handling, not by vanishing gradients.

LSTMs as a partial fix

LSTMs added gates and a cell state with an additive update, which carried gradients further than plain RNNs — an early form of the same "additive path" idea residuals use.

A real-life example

A team trains a small transformer from scratch for Hindi–English code-mixed intent detection, with 24 layers. They copy an old configuration that places LayerNorm after the residual add (post-norm) and use no warmup. Training loss stays flat at the starting value for 5,000 steps. Per-layer gradient norms show the bottom 8 layers receiving about a thousandth of the gradient of the top layers.

Switching to pre-norm and adding 1,000 warmup steps gets the loss falling within the first few hundred steps. The model itself did not change; the path the gradient could travel did.

Follow-up questions to expect

  • "How do you detect vanishing gradients?" — Log gradient norms per layer; early layers near zero while late layers are normal is the signature.
  • "What is pre-norm versus post-norm?" — Pre-norm normalises the input to each sub-layer and leaves the residual path clean; post-norm normalises after the add. Pre-norm is easier to train deep; almost all modern LLMs use it.
  • "Why did ReLU help?" — Its slope is exactly 1 for positive inputs, so it does not shrink gradients the way sigmoid does.