Course Content
Transformer Architecture Q&A
6 sections · 60 lessons
Compare GPT-2 sizes—how do depth, width, and head count scale across models?
What you need to know
| Model | Layers | d_model | Heads | d_head | Params |
|---|---|---|---|---|---|
| GPT-2 small | 12 | 768 | 12 | 64 | 124M |
| GPT-2 medium | 24 | 1,024 | 16 | 64 | 355M |
| GPT-2 large | 36 | 1,280 | 20 | 64 | 774M |
| GPT-2 XL | 48 | 1,600 | 25 | 64 | 1,558M |
Where the parameters come from
Each block has about 12·d² weights: 4·d² in attention (Q, K, V, O) and 8·d² in the MLP (d → 4d → d).
1V, T = 50257, 10242for name, L, d in [("small", 12, 768), ("medium", 24, 1024), ("large", 36, 1280), ("xl", 48, 1600)]:3 blocks = L * (12 * d * d + 13 * d) # 12·d² weights + biases/norms4 total = V * d + T * d + blocks + 2 * d5 print(f"{name:6} {total/1e6:7.0f}M blocks {blocks/total:.0%} embeddings {V*d/total:.0%}")6# small 124M blocks 68% embeddings 31%7# medium 355M blocks 85% embeddings 15%8# large 774M blocks 92% embeddings 8%9# xl 1558M blocks 95% embeddings 5%Three patterns to name
d_headis fixed at 64. Head count isd_model / 64, so width and head count grow together by design.- Depth grows faster than width. Depth rises 4× and width about 2× from small to XL. Since blocks cost
L·d², that gives about4 × 2.1² ≈ 17×more block parameters. - Embeddings are a fixed cost.
V × dgrows only with width, so bigger models spend almost everything on blocks. That is why scaling studies report non-embedding parameters.
Does the shape matter?
Kaplan et al. (2020), "Scaling Laws for Neural Language Models", found that at a fixed non-embedding parameter count, loss depends only weakly on the exact depth-to-width ratio over a wide range. Size and data matter far more than shape.
But shape matters for latency. Layers run one after another, while width is spread across GPU cores. At batch size 1, a deeper, narrower model has more sequential steps per token and is often slower than a shallower, wider one with the same parameters.
GPT-3 continued the recipe to 96 layers, d_model 12,288 and 96 heads of 128.
A real-life example
A code-completion assistant must run on developers' laptops with a 50 ms per-token budget. The team compares two made-up configurations of about 350M parameters: 24 layers at width 1,024 (GPT-2 medium's shape), and 12 layers at width about 1,450.
Their test shows almost the same completion accuracy, as Kaplan's result would suggest. But on the laptop GPU, the 12-layer model is noticeably faster per token, because each token passes through half as many sequential layers and each layer's matmuls are wide enough to keep the GPU busy. They ship the shallower model. On a server at large batch sizes, the difference would be much smaller.
Follow-up questions to expect
- "Why is
d_headkept at 64 rather than adding width per head?" — It keeps each head's dot product well-conditioned and kernels efficient, and adds more heads (more attention patterns) as width grows. - "What is
12·d²made of?" —4·d²for Q, K, V and O, and8·d²for the two MLP matrices ofd × 4d. - "Why do larger models look 'cheaper' in embedding terms?" —
V·dgrows only linearly in width, while blocks grow withL·d².