Transformer Architecture Q&A

Course Content

Transformer Architecture Q&A

6 sections · 60 lessons

Explain role of residual connections in a Transformer block and what fails without them?


A pre-norm block on the residual streamResidual stream x, width d_modelx = x + Attn(RMSNorm(x))x = x + FFN(RMSNorm(x))Repeat for every layer; identity path untouchedFinal norm, then the LM head
In the same 48-layer stack the input gradient is about 2e-11 without skips and arrives intact with them — the identity path is what makes depth trainable.

What you need to know

The gradient argument

The derivative of y = x + f(x) is:

Text
dy/dx = I + f'(x)

The I (identity) means that even if f'(x) is small, the gradient passes through at full strength. Without the skip, each layer multiplies the gradient by f'(x). If that shrinks it to 0.5, then after 32 layers the gradient is 0.5^32 ≈ 2 × 10^-10 — effectively zero.

We checked this with 48 random tanh layers of width 64:

Python
import torchtorch.manual_seed(0)d, depth = 64, 48layers = [torch.nn.Linear(d, d) for _ in range(depth)]def run(x, residual):    for f in layers:        x = x + torch.tanh(f(x)) if residual else torch.tanh(f(x))    return xfor residual in (False, True):    x = torch.randn(8, d, requires_grad=True)    run(x, residual).sum().backward()    print(f"residual={residual}: grad norm at input = {x.grad.norm():.2e}")# residual=False: grad norm at input = 2.24e-11# residual=True: grad norm at input = 2.50e+02

Same layers, same input. Without skips, the gradient at the input is about 10^-11, so the early layers never learn. With skips it arrives intact.

The residual stream

Think of the residual as a bus of width d_model running through the whole model. Each attention and FFN sub-layer reads from it and adds a correction. Layers make small edits instead of rebuilding the representation, and a layer that is not useful for an input can contribute almost nothing. Interpretability work describes most model behaviour in these terms.

Where the norm sits: pre-norm versus post-norm

Post-norm (2017 original)

  • x = Norm(x + f(x))
  • Norm sits on the residual path
  • Needs careful learning-rate warm-up
  • Unstable when very deep

Pre-norm (modern default)

  • x = x + f(Norm(x))
  • Residual path is a clean identity
  • Trains stably at 80+ layers
  • Needs one final norm before the output head

Pre-norm keeps the identity path untouched, which is why GPT-2 onwards, Llama, Qwen and most 2026 LLMs use it.

A real-life example

A team trains a 48-layer code-completion model. To try an idea, an engineer writes a custom block with the attention output passed straight into the FFN, dropping the first skip connection. Training loss falls for the first few hundred steps, then stalls near the level of a model that predicts only common tokens like ) and newline. Gradient-norm logs show the bottom 20 layers receiving almost no gradient.

Restoring x = x + attn(norm(x)) fixed training immediately. A second lesson came from an older post-norm version of the same model: it diverged unless the warm-up was long. Switching to pre-norm let them shorten the warm-up and raise the learning rate.

Follow-up questions to expect

  • "Why does pre-norm need a final norm?" — The residual stream is never normalised inside the blocks, so its scale grows with depth; one last norm fixes it before the output head.
  • "Does the residual limit what a layer can do?" — No. The layer can still output any update; it just does not have to relearn the identity.
  • "Where else are residuals used?" — ResNets in vision introduced them; nearly every deep architecture now uses them.