Transformer Architecture Q&A

Course Content

Transformer Architecture Q&A

6 sections · 60 lessons

Compare GPT-2 sizes—how do depth, width, and head count scale across models?


What you need to know

ModelLayersd_modelHeadsd_headParams
GPT-2 small127681264124M
GPT-2 medium241,0241664355M
GPT-2 large361,2802064774M
GPT-2 XL481,60025641,558M

Where the parameters come from

Each block has about 12·d² weights: 4·d² in attention (Q, K, V, O) and 8·d² in the MLP (d → 4d → d).

Python
V, T = 50257, 1024for name, L, d in [("small", 12, 768), ("medium", 24, 1024), ("large", 36, 1280), ("xl", 48, 1600)]:    blocks = L * (12 * d * d + 13 * d)          # 12·d² weights + biases/norms    total = V * d + T * d + blocks + 2 * d    print(f"{name:6} {total/1e6:7.0f}M  blocks {blocks/total:.0%}  embeddings {V*d/total:.0%}")# small      124M  blocks 68%  embeddings 31%# medium     355M  blocks 85%  embeddings 15%# large      774M  blocks 92%  embeddings 8%# xl        1558M  blocks 95%  embeddings 5%

Three patterns to name

  • d_head is fixed at 64. Head count is d_model / 64, so width and head count grow together by design.
  • Depth grows faster than width. Depth rises 4× and width about 2× from small to XL. Since blocks cost L·d², that gives about 4 × 2.1² ≈ 17× more block parameters.
  • Embeddings are a fixed cost. V × d grows only with width, so bigger models spend almost everything on blocks. That is why scaling studies report non-embedding parameters.

Does the shape matter?

Kaplan et al. (2020), "Scaling Laws for Neural Language Models", found that at a fixed non-embedding parameter count, loss depends only weakly on the exact depth-to-width ratio over a wide range. Size and data matter far more than shape.

But shape matters for latency. Layers run one after another, while width is spread across GPU cores. At batch size 1, a deeper, narrower model has more sequential steps per token and is often slower than a shallower, wider one with the same parameters.

GPT-3 continued the recipe to 96 layers, d_model 12,288 and 96 heads of 128.

A real-life example

A code-completion assistant must run on developers' laptops with a 50 ms per-token budget. The team compares two made-up configurations of about 350M parameters: 24 layers at width 1,024 (GPT-2 medium's shape), and 12 layers at width about 1,450.

Their test shows almost the same completion accuracy, as Kaplan's result would suggest. But on the laptop GPU, the 12-layer model is noticeably faster per token, because each token passes through half as many sequential layers and each layer's matmuls are wide enough to keep the GPU busy. They ship the shallower model. On a server at large batch sizes, the difference would be much smaller.

Follow-up questions to expect

  • "Why is d_head kept at 64 rather than adding width per head?" — It keeps each head's dot product well-conditioned and kernels efficient, and adds more heads (more attention patterns) as width grows.
  • "What is 12·d² made of?" — 4·d² for Q, K, V and O, and 8·d² for the two MLP matrices of d × 4d.
  • "Why do larger models look 'cheaper' in embedding terms?" — V·d grows only linearly in width, while blocks grow with L·d².