Course Content
How Large Language Models Work
3 sections · 9 lessons
Mini-Project: Analyze a Real LLM
Reading about tokenisers, attention heads and sampling gets you a working mental model. Opening a real model and measuring those things gets you something better: calibration. You find out that attention maps are messier than the diagrams, that a specific head really does do the thing it was supposed to do, and that a number you assumed was around 0.9 is actually 0.4.
This project takes GPT-2 small apart with about a hundred lines of code. It is a 124-million-parameter model from 2019 — old, weak by current standards, and perfect for this, because it runs on a laptop CPU in seconds and its internals are identical in kind to those of a model a thousand times larger. Every measurement here transfers.
Work through it in a notebook and actually run the code. The point is the numbers you get, not the numbers written here.
Setting up
pip install torch transformers matplotlib numpy1import torch2import numpy as np3from transformers import AutoTokenizer, AutoModelForCausalLM45tok = AutoTokenizer.from_pretrained("gpt2")6model = AutoModelForCausalLM.from_pretrained(7 "gpt2", attn_implementation="eager" # required to read attention weights8).eval()910cfg = model.config11print(f"layers {cfg.n_layer}") # 1212print(f"heads/layer {cfg.n_head}") # 1213print(f"d_model {cfg.n_embd}") # 76814print(f"vocab {cfg.vocab_size}") # 5025715print(f"max ctx {cfg.n_positions}") # 102416print(f"params {sum(p.numel() for p in model.parameters()):,}")Before running anything else, do one calculation by hand from those numbers and check it against the printed parameter count.
Embedding table: 50,257 x 768 = 38,597,376 (31% of the model!)Position table: 1,024 x 768 = 786,432Per layer: attention QKVO: 4 x 768 x 768 = 2,359,296 FFN (4x wide): 2 x 768 x 3072 = 4,718,592 --------- 7,077,888 x 12 layers = 84,934,656Total (weights, ignoring biases and norms): about 124 million.Two facts are already visible. Almost a third of this model is the token lookup table — small models spend a disproportionate share of themselves on vocabulary. And within the transformer stack, the feed-forward layers hold twice as many parameters as attention does.
Task 1 — Tokenisation forensics
Measure how the tokeniser treats different kinds of text, and find where it fragments.
1samples = {2 "english prose": "The committee reached a decision on Thursday afternoon.",3 "python code": "def f(x):\n return [i**2 for i in range(x)]",4 "json": '{"user_id": 4471, "status": "active", "score": 0.87}',5 "numbers": "The result was 1234567.89 in Q3 2024.",6 "rare words": "The pharmacokinetics of levothyroxine remain idiosyncratic.",7 "identifier": "order-8f14e45fceea167a5a36dedd4bea2543",8 "german": "Donaudampfschifffahrtsgesellschaft ist ein langes Wort.",9}1011for name, text in samples.items():12 ids = tok.encode(text)13 print(f"{name:14s} {len(text):3d} chars {len(ids):3d} tokens "14 f"{len(text)/len(ids):4.1f} chars/token")Then look at where the splits fall for the worst offender:
worst = "order-8f14e45fceea167a5a36dedd4bea2543"pieces = tok.convert_ids_to_tokens(tok.encode(worst))print(len(pieces), pieces)Two specific things to check while you are here, because both cause real bugs in production systems:
1# 1. A leading space is part of the token, not a separator2print(tok.encode("hello")) # [31373]3print(tok.encode(" hello")) # [23748] - a completely different token45# 2. Trailing whitespace pushes the model off a normal boundary6print(tok.convert_ids_to_tokens(tok.encode("The capital of France is")))7print(tok.convert_ids_to_tokens(tok.encode("The capital of France is ")))What you should find: the prose sentence lands above the usual 4-characters-per-token rule of thumb (about 6 here), because every word is common and each token carries its leading space; code and JSON drop well below that because of indentation, brackets and quoting; the hex identifier fragments into a dozen or more pieces; the long German compound splits into fragments that do not respect its actual morpheme boundaries. And the trailing-space version ends with a lone space token that changes the distribution over what can come next.
Task 2 — Surprisal, perplexity and entropy
Now measure what the model actually believes. Two different quantities are worth separating carefully:
- Surprisal −logp(actual next token) — how surprised the model was by what really came next. Needs the true text.
- Entropy −∑ipilogpi — how spread out the model's distribution was. Needs no ground truth; measurable during generation.
1def analyse(text):2 ids = tok(text, return_tensors="pt").input_ids3 with torch.no_grad():4 logits = model(ids).logits56 logprobs = torch.log_softmax(logits[0, :-1], dim=-1) # predict positions 1..n-17 targets = ids[0, 1:]89 surprisal = -logprobs[torch.arange(len(targets)), targets] # per token, nats10 probs = logprobs.exp()11 entropy = -(probs * logprobs).sum(-1)1213 tokens = tok.convert_ids_to_tokens(targets)14 return tokens, surprisal, entropy1516text = ("The capital of France is Paris. The capital of Australia is Canberra. "17 "The capital of Mongolia is Ulaanbaatar.")18tokens, surprisal, entropy = analyse(text)1920print(f"mean surprisal {surprisal.mean():.3f} nats")21print(f"perplexity {surprisal.mean().exp():.1f}")22print()23print(f"{'token':>14s} {'surprisal':>10s} {'entropy':>8s}")24for t, s, e in zip(tokens, surprisal, entropy):25 print(f"{t:>14s} {s:10.3f} {e:8.3f}")Read the per-token column rather than the average. The interesting rows are the extremes:
| Pattern | Surprisal | Entropy | Interpretation |
|---|---|---|---|
| Function words, second half of a repeated phrase | Very low | Low | The model is confident and correct — structure it has fully absorbed |
| The token " Paris" after "capital of France is" | Low | Low | A fact it holds firmly |
| The first piece of a rare proper noun | High | High | Genuine uncertainty — many continuations were plausible |
| The later pieces of that same noun | Low | Low | Once committed to a multi-token word, finishing it is nearly forced |
| High surprisal, low entropy | High | Low | The model was confident and wrong. The most diagnostic combination there is. |
That last row is worth hunting for deliberately, because it is what a confabulation looks like from the inside. A low-entropy distribution means the model saw one obvious answer; a high surprisal on the real token means the obvious answer was not the true one.
To see the branching-factor reading of perplexity directly, compare two texts:
1for label, s in [2 ("boilerplate", "The terms and conditions of this agreement shall be "3 "governed by and construed in accordance with the laws of"),4 ("surprising", "The quantum eigenstate collapsed into a lukewarm bowl "5 "of Tuesday afternoon regret"),6]:7 _, sur, _ = analyse(s)8 print(f"{label:12s} perplexity {sur.mean().exp():7.1f}")The gap should be large — legal boilerplate is nearly deterministic, deliberate nonsense is not. Note also that a perplexity of k means the model was on average as uncertain as if choosing uniformly among k options, out of a vocabulary of 50,257.
Task 3 — Finding an induction head
This is the most rewarding part of the project, because you are going to locate a specific, named circuit inside a specific head.
An induction head implements pattern completion: having seen [A][B] earlier in the context, when it encounters [A] again it attends to the token that followed the first [A] — that is, to [B] — so the model can predict [B]. It is a large part of why putting examples in a prompt works at all.
The test is elegant: feed the model a sequence of random tokens repeated twice. There is no semantic content, no grammar, nothing to memorise. The only usable signal is the repetition itself. If a head attends from position i in the second copy back to position i−n+1 in the first copy, it is doing induction and nothing else.
1torch.manual_seed(0)2n = 403half = torch.randint(1000, 20000, (n,))4ids = torch.cat([half, half]).unsqueeze(0) # random sequence, twice56with torch.no_grad():7 out = model(ids, output_attentions=True)89scores = {}10for layer, attn in enumerate(out.attentions): # (1, heads, seq, seq)11 for head in range(attn.shape[1]):12 a = attn[0, head]13 # for each query in the second copy, the induction target is i - n + 114 vals = [a[i, i - n + 1].item() for i in range(n, 2 * n)]15 scores[(layer, head)] = float(np.mean(vals))1617for (layer, head), score in sorted(scores.items(), key=lambda kv: -kv[1])[:6]:18 print(f"layer {layer:2d} head {head:2d} induction score {score:.3f}")Random chance for this measurement is roughly 1/n, about 0.025. Strong induction heads score an order of magnitude above that. In GPT-2 small they show up in the middle layers rather than the first or last — which makes mechanistic sense, since an induction head needs a previous-token head in an earlier layer to have already written "the token before me was X" into the residual stream before it can match on it.
While the attention tensors are in hand, measure the attention sink:
1real = tok("The trophy would not fit in the brown suitcase because it was "2 "too small.", return_tensors="pt")3with torch.no_grad():4 out = model(**real, output_attentions=True)56for layer, attn in enumerate(out.attentions):7 to_first = attn[0, :, 1:, 0].mean().item() # all heads, queries after 0, key 08 print(f"layer {layer:2d} mean attention on token 0: {to_first:.3f}")You will find layers where heads dump a large share of their weight on the very first token regardless of content. This is not a bug. Softmax must sum to 1, so a head with nothing relevant to retrieve needs somewhere to park its mass, and the first token becomes that null option. It is also why long-context serving systems deliberately pin the first few tokens in the cache — evict them and output quality falls far more than their content would suggest.
Task 4 — Decoding strategies, measured
Same model, same prompt, different decoders. Quantify the difference instead of eyeballing it.
1def distinct_n(text, n=2):2 words = text.split()3 if len(words) < n:4 return 0.05 grams = [tuple(words[i:i+n]) for i in range(len(words) - n + 1)]6 return len(set(grams)) / len(grams)78def longest_repeat(text, n=4):9 """Length of the longest run of repeated 4-grams - a loop detector."""10 words = text.split()11 grams = [tuple(words[i:i+n]) for i in range(len(words) - n + 1)]12 seen, best, run = set(), 0, 013 for g in grams:14 run = run + 1 if g in seen else 015 seen.add(g)16 best = max(best, run)17 return best1819prompt = tok("The future of artificial intelligence", return_tensors="pt")2021configs = {22 "greedy": dict(do_sample=False),23 "T=0.5": dict(do_sample=True, temperature=0.5, top_k=0),24 "T=1.0": dict(do_sample=True, temperature=1.0, top_k=0),25 "T=1.5": dict(do_sample=True, temperature=1.5, top_k=0),26 "top-k 50": dict(do_sample=True, temperature=1.0, top_k=50),27 "top-p 0.9": dict(do_sample=True, temperature=1.0, top_p=0.9, top_k=0),28 "beam 5": dict(do_sample=False, num_beams=5, early_stopping=True),29}3031torch.manual_seed(0)32for name, kw in configs.items():33 out = model.generate(**prompt, max_new_tokens=80,34 pad_token_id=tok.eos_token_id, **kw)35 text = tok.decode(out[0], skip_special_tokens=True)36 print(f"{name:10s} distinct-2 {distinct_n(text):.2f} "37 f"loop-run {longest_repeat(text):2d}")38 print(" ", text[len(tok.decode(prompt.input_ids[0])):][:110].replace("\n", " "))39 print()What to look for in the two numbers:
| Setting | Expected distinct-2 | Expected loop-run | Characteristic output |
|---|---|---|---|
| Greedy | Low | High | Often falls into a repeating phrase and never escapes |
| Beam search, width 5 | Lowest | Highest | Even more repetitive — it is a better probability maximiser, and maximal probability is repetitive |
| T = 0.5 | Moderate | Low | Safe, coherent, dull |
| T = 1.0 | High | 0 | Varied and mostly coherent |
| T = 1.5 | Highest | 0 | Coherence breaks down; rare tokens leak in |
| Top-p 0.9 | High | 0 | Usually the best balance — the tail is cut before sampling |
The beam search row is the one worth sitting with. Beam search is strictly better than greedy at finding high-probability sequences, and it produces worse open-ended text. That is not a contradiction; it is evidence that human writing does not sit at the mode of the model's distribution.
Task 5 — The memory arithmetic
One last measurement, done with a calculator rather than code, because it is the number that decides deployments.
GPT-2 small KV cache, per token: 2 (K and V) x 12 heads x 64 head_dim = 1,536 values per layer x 12 layers = 18,432 values x 4 bytes (fp32) = 73,728 bytes = 72 KiB per tokenA full 1,024-token context: 73,728 x 1,024 = 75,497,472 bytes = 75.5 MB (72 MiB) - for one sequence.Verify it against the real tensors:
1ids = torch.randint(0, 50000, (1, 1024))2with torch.no_grad():3 out = model(ids, use_cache=True)45# Each cached layer holds a key tensor and a value tensor.6# (Indexing rather than unpacking works on older and newer transformers.)7total = sum(layer[0].numel() + layer[1].numel() for layer in out.past_key_values)8print(f"{len(out.past_key_values)} layers, "9 f"{total:,} cached values, {total * 4 / 1e6:.1f} MB in fp32")Now scale the same formula to a 70-billion-parameter model with 80 layers, 64 heads and head dimension 128 in fp16, if every head kept its own keys and values: 2.62 MB per token, 10.7 GB for a 4,096-token conversation, per concurrent user. That single number is why grouped-query attention exists, and why a self-hosted deployment usually runs out of memory long before it runs out of arithmetic.
Turning measurements into judgement
Once the code runs, push it somewhere it will surprise you. Run the tokenisation audit over a real sample of your own data rather than the toy strings above, and see whether your chars-per-token assumption survives. Run the surprisal analysis over a document the model gets subtly wrong, and check whether the incorrect tokens show the confident-and-wrong signature — high surprisal with low entropy — because if they do, entropy becomes a usable runtime signal for flagging outputs that need checking.
Compare a base model against its instruction-tuned sibling on identical inputs. The base model's entropy on the token right after a question mark will typically be much higher; the tuned model has learned that an answer, not another question, comes next. That difference is instruction tuning, visible as a number.
And keep the induction-head experiment. It is a clean demonstration that these models are not undifferentiated soup: there are identifiable components doing identifiable jobs, findable with forty random tokens and a mean over one diagonal of an attention matrix. Everything else you will read about interpretability builds on measurements of exactly that shape.