Course Content
Transformer Architecture Q&A
6 sections · 60 lessons
When visualizing attention weights, what patterns appear and what do they indicate?
What you need to know
The patterns and what they mean
| Pattern | Looks like | What it suggests |
|---|---|---|
| Local band | bright diagonal and just below | local context, syntax |
| Previous-token | one line exactly one step below the diagonal | building block for induction |
| Attention sink | bright first column | a place to dump probability mass |
| Induction | attends to the token after an earlier copy of the current token | copying, in-context learning |
| Delimiter / structure | focus on commas, brackets, quotes, newlines | tracking structure |
| Broad / uniform | even spread | averaging, bag-of-words features |
Attention sinks
Softmax weights must sum to 1, so a head always attends somewhere, even when no token is relevant. Models learn to park that weight on the first token. Xiao et al. (2023), in StreamingLLM, showed that a sliding-window cache works well only if it keeps these first few "sink" tokens. OpenAI's gpt-oss models (2025) add a learned sink value per head to the softmax denominator, so the head can give low weight to all real tokens.
Induction heads
Olsson et al. (Anthropic, 2022) described induction heads: a previous-token head in an earlier layer writes "the token before me was A" into each position; a later head at a new occurrence of A looks for the position whose previous token was A, and copies what came next. Given "... Mr Sharma ... Mr", it predicts "Sharma". They linked the formation of these heads to a jump in in-context learning during training.
How to look
transformersmodels acceptoutput_attentions=True; this needs the "eager" attention implementation, because fused kernels never build the matrix.- Tools such as BertViz and TransformerLens draw per-head heatmaps and support ablation.
The caveat
A weight tells you how much of a value vector is mixed in, not how large or useful that value is. A head can put 90% of its weight on a sink whose value vector is near zero; its real contribution comes from the other 10%. Value norms, W_o, and the residual stream all matter. Test a hypothesis by removing the head (ablation) or replacing its activations (activation patching) and measuring the change in output.
A real-life example
A code-completion assistant keeps suggesting the wrong variable name inside long functions. An engineer visualises the heads at the position where the model should write customer_id. In one later-layer head, the weight sits on the token right after an earlier customer_ — the induction pattern — but the earlier occurrence it picks is customer_name, from a different function 3,000 tokens back.
They ablate that head and the error rate on a test set of 200 such completions barely changes, so it is not the only cause. They then find the model's retrieval of the right function's scope is weak beyond 2,000 tokens. The fix is on the input side: the completion service now puts the current function's text last, closest to the cursor. The heatmap gave the hypothesis; the ablation kept them from "fixing" the wrong thing.
Follow-up questions to expect
- "Why do so many heads attend to the first token?" — It is an attention sink: a learned place to put weight when nothing else is relevant, because softmax cannot output all zeros.
- "What is an induction head?" — A pair of heads that together copy the token that followed an earlier occurrence of the current token; a key mechanism for in-context learning.
- "How do you know a head matters?" — Ablate or patch it and measure the change in loss or behaviour; the weight map alone does not show importance.