Transformer Architecture Q&A

Course Content

Transformer Architecture Q&A

6 sections · 60 lessons

Architecturally, how does GPT-3 differ from GPT-2 beyond just scale?


What you need to know

The real differences

GPT-2 XLGPT-3 175B
Attentiondense in every layeralternating dense and locally banded sparse
Context1,0242,048
Layers4896
d_model1,60012,288
Heads × d_head25 × 6496 × 128
Parameters1.5B175B
Training tokensWebText (about 40 GB of text)about 300B tokens, mixed sources

The block, the pre-norm layout, the GELU MLP, learned position embeddings and the byte-level BPE vocabulary of 50,257 all stayed the same.

Checking the size with the formulas

Text
blocks:     12 × L × d² = 12 × 96 × 12,288²  ≈ 173.9Bembeddings: 50,257 × 12,288                   ≈   0.6Btotal                                          ≈ 174.6B  → "175B"training compute ≈ 6 × N × D = 6 × 175e9 × 300e9 ≈ 3.15e23 FLOPs

The GPT-3 paper reports about 3.14e23 FLOPs for training, which matches the rule of thumb: 2 FLOPs per parameter per token for the forward pass and 4 for the backward.

What GPT-3 actually showed

  • Few-shot in-context learning. Put a few examples in the prompt and the model does the task, with no weight updates. This improved steadily with model size in the paper's experiments.
  • A family of sizes. Eight models from 125M to 175B with the same recipe, so the effect of scale could be seen directly.

What came after

Hoffmann et al. (2022), the "Chinchilla" paper, found that models like GPT-3 were under-trained for their size: for a fixed compute budget, parameters and tokens should grow together, roughly 20 tokens per parameter. GPT-3 used under 2 tokens per parameter. Modern models go much further — Llama 3 8B was trained on over 15T tokens, nearly 2,000 per parameter — because a smaller, longer-trained model is cheaper to serve.

A real-life example

An e-commerce team needs to sort new product listings into 400 categories. In the GPT-2 era, they would have fine-tuned a model on thousands of labelled listings per category group. The GPT-3 result is what lets them start differently in 2026: give a large model the category list and ten labelled examples in the prompt, and it classifies new listings with no training at all.

They use that few-shot setup to launch in a week, measure it on 2,000 hand-labelled listings, and later fine-tune a small model once they have collected enough labels to cut the per-listing cost. The architecture of the model they call is not the point; the capability that scale gave it is.

Follow-up questions to expect

  • "Why the sparse layers?" — To make 2,048-token attention affordable at 96 layers; half the layers attend only to a local band.
  • "Was GPT-3 compute-optimal?" — By Chinchilla's later analysis, no: it had too many parameters for its 300B tokens.
  • "What is in-context learning mechanically?" — The model conditions on the examples through attention; no weights change. Induction-style heads are one mechanism that supports it.