Transformer Architecture Q&A

Course Content

Transformer Architecture Q&A

6 sections · 60 lessons

Contrast 2017 Transformer block vs Llama-3-style block—what concrete changes?


What you need to know

2017 TransformerLlama-3-style
Norm placementpost-norm: LN(x + Sublayer(x))pre-norm: x + Sublayer(Norm(x))
Norm typeLayerNorm (mean, variance, bias)RMSNorm (scale only)
FFNReLU, d → 4d → d, 2 matricesSwiGLU, d → ~8/3·d → d, 3 matrices
Positionsinusoidal, added to inputRoPE on Q and K, every layer
AttentionMHA (n_kv = n_q)GQA (32 query, 8 KV heads in 8B)
Biasesyesnone
Dropout0.10
Text
2017:   h = LN(x + Attn(x))            y = LN(h + FFN_ReLU(h))Llama:  h = x + Attn_GQA_RoPE(RMSNorm(x))   y = h + SwiGLU(RMSNorm(h))

The two new pieces, in code

Python
import torch, torch.nn as nn, torch.nn.functional as Fclass RMSNorm(nn.Module):    def __init__(self, d, eps=1e-5):        super().__init__()        self.eps, self.weight = eps, nn.Parameter(torch.ones(d))    def forward(self, x):   # no mean subtraction, no bias        return self.weight * x * torch.rsqrt(x.pow(2).mean(-1, keepdim=True) + self.eps)class SwiGLU(nn.Module):    def __init__(self, d, hidden):        super().__init__()        self.gate = nn.Linear(d, hidden, bias=False)        self.up = nn.Linear(d, hidden, bias=False)        self.down = nn.Linear(hidden, d, bias=False)    def forward(self, x):        return self.down(F.silu(self.gate(x)) * self.up(x))x = torch.randn(2, 5, 4096)print(SwiGLU(4096, 14336)(RMSNorm(4096)(x)).shape)   # torch.Size([2, 5, 4096])

Llama-3-8B uses a hidden width of 14,336 (3.5·d, rounded up from 8/3·d for hardware), so its MLP has 3 × 4096 × 14336 ≈ 176M parameters per layer.

Why each change

  • Pre-norm — the residual path is a clean identity, so gradients reach early layers undistorted; 32–126-layer stacks train stably.
  • RMSNorm — drops the mean subtraction and bias; cheaper, with no measurable quality loss (Zhang and Sennrich, 2019).
  • SwiGLU — a gated MLP that gave better quality at matched parameters in Shazeer's 2020 "GLU Variants Improve Transformer" experiments.
  • RoPE — encodes relative position in the Q·K score and can be stretched to long contexts (Su et al., 2021).
  • GQA — 4× smaller KV cache in the 8B model, critical for long-context serving.
  • No biases, no dropout — biases add little at scale; dropout is unnecessary for single-pass training on trillions of tokens.

What changed again after Llama 3 (2024–2026)

  • QK-norm on queries and keys for stable training (OLMo 2, Qwen3, Gemma 3).
  • MoE FFNs in place of the dense SwiGLU (DeepSeek-V3, Qwen3's MoE models, Llama 4, gpt-oss).
  • Multi-head latent attention to compress the KV cache (DeepSeek-V2 and V3).
  • Interleaved local and global attention layers (Gemma 2 and 3, gpt-oss).

A real-life example

A code-completion team has an in-house 350M GPT-2-style model and wants to modernise it before training a longer-context version. They change one part at a time and measure each on their held-out code set: pre-norm with RMSNorm first (training no longer diverges at a higher learning rate), then SwiGLU with a narrower hidden width to keep the parameter count equal, then RoPE so they can extend context from 2K to 16K with a short fine-tune, then GQA with 4 KV heads so the IDE backend can hold more open files in the KV cache.

Changing one component at a time let them attribute each gain. When they later loaded Llama-family weights into their own serving code, their parity test caught a missed detail: RMSNorm's epsilon (1e-5 in Llama 3's config) and the RoPE base frequency (500,000 in Llama 3, versus 10,000 in the original RoPE) must match the checkpoint exactly.

Follow-up questions to expect

  • "Why is the SwiGLU width 8/3·d and not 4d?" — It has three matrices instead of two; 3 × 8/3 = 8, the same parameter count as a 4d two-matrix MLP.
  • "Why does pre-norm need a final norm?" — The residual stream is never normalised inside the blocks, so its scale grows with depth; a final RMSNorm fixes that before the LM head.
  • "Is post-norm ever better?" — It can give slightly better final quality when it trains stably, but it needs careful warmup and fails more often at depth; most large models use pre-norm.