Course Content
Transformer Architecture Q&A
6 sections · 60 lessons
Define weight tying in GPT-2 and why it’s 124M params vs ~163M raw?
What you need to know
The count, verified
1import torch.nn as nn2V, T_max, d, L = 50257, 1024, 768, 123wte, wpe, ln_f = nn.Embedding(V, d), nn.Embedding(T_max, d), nn.LayerNorm(d)4# nn.TransformerEncoderLayer has the same weights as a GPT-2 block:5# fused QKV + bias, W_o + bias, 768→3072→768 MLP + biases, two LayerNorms6block = nn.TransformerEncoderLayer(d, nhead=12, dim_feedforward=4 * d, norm_first=True)78count = lambda m: sum(p.numel() for p in m.parameters())9tied = count(wte) + count(wpe) + L * count(block) + count(ln_f)10print(f"{count(block):,} per block") # 7,087,87211print(f"{tied:,} with a tied LM head") # 124,439,80812print(f"{tied + V * d:,} with a separate LM head") # 163,037,184token embedding 50,257 × 768 = 38.6Mpositions 1,024 × 768 = 0.8M12 blocks × 7.09M = 85.1M (about 12·d² each, plus biases and norms)final norm ≈ 0.0Mtotal (tied) = 124.4M+ separate LM head = 163.0MWhy it works
- Symmetry of the job. The input embedding says "this token's meaning is this vector"; the output head says "this vector means this token is likely". Sharing forces the two to agree.
- More training signal. Rare tokens appear rarely as inputs, but the output softmax updates every row at every step, so tied embeddings of rare tokens are trained more.
- Evidence. Press and Wolf (2017), "Using the Output Embedding to Improve Language Models", showed that tying lowers perplexity and saves parameters in language models.
When models untie
Untying gives the model more freedom: input and output roles can differ. At large sizes the extra V × d matrix is a small share, so it is affordable:
| Model | Vocab × d | Tied? | Embedding share of total |
|---|---|---|---|
| GPT-2 small | 50,257 × 768 | yes | ~31% |
| Llama 3.2 1B | 128,256 × 2,048 | yes | ~21% (counted once) |
| Llama-3-8B | 128,256 × 4,096 | no | ~13% (two matrices) |
Small models with large vocabularies tie, because an untied head would be a big fraction of the model.
A real-life example
A team builds a 135M-parameter model for on-device next-word suggestions in a Hindi-English keyboard app. They use a 64,000-token vocabulary with d = 768, so each embedding matrix is 49M parameters. Untied, the head alone would add 49M parameters — about 36% more to download and hold in phone memory.
They tie the weights. The app's download stays small, and in their tests the tied model's quality matches the untied one, because at this size the two roles benefit from sharing. For a later 8B server model, the same team leaves them untied, since the extra matrix is a small share there.
Follow-up questions to expect
- "Is anything removed at inference?" — No. Both names point to one tensor; the computation is the same, only memory is saved.
- "Does tying constrain the model?" — Yes, input and output embeddings must be the same; large models untie to relax that.
- "How do you tie in PyTorch?" — Assign the same parameter:
lm_head.weight = wte.weight, so gradients from both uses add into one tensor.