Course Content
Transformer Architecture Q&A
6 sections · 60 lessons
Contrast 2017 Transformer block vs Llama-3-style block—what concrete changes?
What you need to know
| 2017 Transformer | Llama-3-style | |
|---|---|---|
| Norm placement | post-norm: LN(x + Sublayer(x)) | pre-norm: x + Sublayer(Norm(x)) |
| Norm type | LayerNorm (mean, variance, bias) | RMSNorm (scale only) |
| FFN | ReLU, d → 4d → d, 2 matrices | SwiGLU, d → ~8/3·d → d, 3 matrices |
| Position | sinusoidal, added to input | RoPE on Q and K, every layer |
| Attention | MHA (n_kv = n_q) | GQA (32 query, 8 KV heads in 8B) |
| Biases | yes | none |
| Dropout | 0.1 | 0 |
2017: h = LN(x + Attn(x)) y = LN(h + FFN_ReLU(h))Llama: h = x + Attn_GQA_RoPE(RMSNorm(x)) y = h + SwiGLU(RMSNorm(h))The two new pieces, in code
1import torch, torch.nn as nn, torch.nn.functional as F23class RMSNorm(nn.Module):4 def __init__(self, d, eps=1e-5):5 super().__init__()6 self.eps, self.weight = eps, nn.Parameter(torch.ones(d))7 def forward(self, x): # no mean subtraction, no bias8 return self.weight * x * torch.rsqrt(x.pow(2).mean(-1, keepdim=True) + self.eps)910class SwiGLU(nn.Module):11 def __init__(self, d, hidden):12 super().__init__()13 self.gate = nn.Linear(d, hidden, bias=False)14 self.up = nn.Linear(d, hidden, bias=False)15 self.down = nn.Linear(hidden, d, bias=False)16 def forward(self, x):17 return self.down(F.silu(self.gate(x)) * self.up(x))1819x = torch.randn(2, 5, 4096)20print(SwiGLU(4096, 14336)(RMSNorm(4096)(x)).shape) # torch.Size([2, 5, 4096])Llama-3-8B uses a hidden width of 14,336 (3.5·d, rounded up from 8/3·d for hardware), so its MLP has 3 × 4096 × 14336 ≈ 176M parameters per layer.
Why each change
- Pre-norm — the residual path is a clean identity, so gradients reach early layers undistorted; 32–126-layer stacks train stably.
- RMSNorm — drops the mean subtraction and bias; cheaper, with no measurable quality loss (Zhang and Sennrich, 2019).
- SwiGLU — a gated MLP that gave better quality at matched parameters in Shazeer's 2020 "GLU Variants Improve Transformer" experiments.
- RoPE — encodes relative position in the Q·K score and can be stretched to long contexts (Su et al., 2021).
- GQA — 4× smaller KV cache in the 8B model, critical for long-context serving.
- No biases, no dropout — biases add little at scale; dropout is unnecessary for single-pass training on trillions of tokens.
What changed again after Llama 3 (2024–2026)
- QK-norm on queries and keys for stable training (OLMo 2, Qwen3, Gemma 3).
- MoE FFNs in place of the dense SwiGLU (DeepSeek-V3, Qwen3's MoE models, Llama 4, gpt-oss).
- Multi-head latent attention to compress the KV cache (DeepSeek-V2 and V3).
- Interleaved local and global attention layers (Gemma 2 and 3, gpt-oss).
A real-life example
A code-completion team has an in-house 350M GPT-2-style model and wants to modernise it before training a longer-context version. They change one part at a time and measure each on their held-out code set: pre-norm with RMSNorm first (training no longer diverges at a higher learning rate), then SwiGLU with a narrower hidden width to keep the parameter count equal, then RoPE so they can extend context from 2K to 16K with a short fine-tune, then GQA with 4 KV heads so the IDE backend can hold more open files in the KV cache.
Changing one component at a time let them attribute each gain. When they later loaded Llama-family weights into their own serving code, their parity test caught a missed detail: RMSNorm's epsilon (1e-5 in Llama 3's config) and the RoPE base frequency (500,000 in Llama 3, versus 10,000 in the original RoPE) must match the checkpoint exactly.
Follow-up questions to expect
- "Why is the SwiGLU width 8/3·d and not 4d?" — It has three matrices instead of two;
3 × 8/3 = 8, the same parameter count as a4dtwo-matrix MLP. - "Why does pre-norm need a final norm?" — The residual stream is never normalised inside the blocks, so its scale grows with depth; a final RMSNorm fixes that before the LM head.
- "Is post-norm ever better?" — It can give slightly better final quality when it trains stably, but it needs careful warmup and fails more often at depth; most large models use pre-norm.