Course Content
Transformer Architecture Q&A
6 sections · 60 lessons
Architecturally, how does GPT-3 differ from GPT-2 beyond just scale?
What you need to know
The real differences
| GPT-2 XL | GPT-3 175B | |
|---|---|---|
| Attention | dense in every layer | alternating dense and locally banded sparse |
| Context | 1,024 | 2,048 |
| Layers | 48 | 96 |
| d_model | 1,600 | 12,288 |
| Heads × d_head | 25 × 64 | 96 × 128 |
| Parameters | 1.5B | 175B |
| Training tokens | WebText (about 40 GB of text) | about 300B tokens, mixed sources |
The block, the pre-norm layout, the GELU MLP, learned position embeddings and the byte-level BPE vocabulary of 50,257 all stayed the same.
Checking the size with the formulas
blocks: 12 × L × d² = 12 × 96 × 12,288² ≈ 173.9Bembeddings: 50,257 × 12,288 ≈ 0.6Btotal ≈ 174.6B → "175B"training compute ≈ 6 × N × D = 6 × 175e9 × 300e9 ≈ 3.15e23 FLOPsThe GPT-3 paper reports about 3.14e23 FLOPs for training, which matches the rule of thumb: 2 FLOPs per parameter per token for the forward pass and 4 for the backward.
What GPT-3 actually showed
- Few-shot in-context learning. Put a few examples in the prompt and the model does the task, with no weight updates. This improved steadily with model size in the paper's experiments.
- A family of sizes. Eight models from 125M to 175B with the same recipe, so the effect of scale could be seen directly.
What came after
Hoffmann et al. (2022), the "Chinchilla" paper, found that models like GPT-3 were under-trained for their size: for a fixed compute budget, parameters and tokens should grow together, roughly 20 tokens per parameter. GPT-3 used under 2 tokens per parameter. Modern models go much further — Llama 3 8B was trained on over 15T tokens, nearly 2,000 per parameter — because a smaller, longer-trained model is cheaper to serve.
A real-life example
An e-commerce team needs to sort new product listings into 400 categories. In the GPT-2 era, they would have fine-tuned a model on thousands of labelled listings per category group. The GPT-3 result is what lets them start differently in 2026: give a large model the category list and ten labelled examples in the prompt, and it classifies new listings with no training at all.
They use that few-shot setup to launch in a week, measure it on 2,000 hand-labelled listings, and later fine-tune a small model once they have collected enough labels to cut the per-listing cost. The architecture of the model they call is not the point; the capability that scale gave it is.
Follow-up questions to expect
- "Why the sparse layers?" — To make 2,048-token attention affordable at 96 layers; half the layers attend only to a local band.
- "Was GPT-3 compute-optimal?" — By Chinchilla's later analysis, no: it had too many parameters for its 300B tokens.
- "What is in-context learning mechanically?" — The model conditions on the examples through attention; no weights change. Induction-style heads are one mechanism that supports it.