Transformer Architecture Q&A

Course Content

Transformer Architecture Q&A

6 sections · 60 lessons

When visualizing attention weights, what patterns appear and what do they indicate?


Patterns that recur in attention heatmapsAttention headLocal band: nearby tokensPrevious-token lineSink: bright first columnInduction: tokenafter last copyDelimiters and bracketsBroad, near-uniform average
Each pattern is a hypothesis about what a head does; only ablating the head shows whether the model relies on it.

What you need to know

The patterns and what they mean

PatternLooks likeWhat it suggests
Local bandbright diagonal and just belowlocal context, syntax
Previous-tokenone line exactly one step below the diagonalbuilding block for induction
Attention sinkbright first columna place to dump probability mass
Inductionattends to the token after an earlier copy of the current tokencopying, in-context learning
Delimiter / structurefocus on commas, brackets, quotes, newlinestracking structure
Broad / uniformeven spreadaveraging, bag-of-words features

Attention sinks

Softmax weights must sum to 1, so a head always attends somewhere, even when no token is relevant. Models learn to park that weight on the first token. Xiao et al. (2023), in StreamingLLM, showed that a sliding-window cache works well only if it keeps these first few "sink" tokens. OpenAI's gpt-oss models (2025) add a learned sink value per head to the softmax denominator, so the head can give low weight to all real tokens.

Induction heads

Olsson et al. (Anthropic, 2022) described induction heads: a previous-token head in an earlier layer writes "the token before me was A" into each position; a later head at a new occurrence of A looks for the position whose previous token was A, and copies what came next. Given "... Mr Sharma ... Mr", it predicts "Sharma". They linked the formation of these heads to a jump in in-context learning during training.

How to look

  • transformers models accept output_attentions=True; this needs the "eager" attention implementation, because fused kernels never build the matrix.
  • Tools such as BertViz and TransformerLens draw per-head heatmaps and support ablation.

The caveat

A weight tells you how much of a value vector is mixed in, not how large or useful that value is. A head can put 90% of its weight on a sink whose value vector is near zero; its real contribution comes from the other 10%. Value norms, W_o, and the residual stream all matter. Test a hypothesis by removing the head (ablation) or replacing its activations (activation patching) and measuring the change in output.

A real-life example

A code-completion assistant keeps suggesting the wrong variable name inside long functions. An engineer visualises the heads at the position where the model should write customer_id. In one later-layer head, the weight sits on the token right after an earlier customer_ — the induction pattern — but the earlier occurrence it picks is customer_name, from a different function 3,000 tokens back.

They ablate that head and the error rate on a test set of 200 such completions barely changes, so it is not the only cause. They then find the model's retrieval of the right function's scope is weak beyond 2,000 tokens. The fix is on the input side: the completion service now puts the current function's text last, closest to the cursor. The heatmap gave the hypothesis; the ablation kept them from "fixing" the wrong thing.

Follow-up questions to expect

  • "Why do so many heads attend to the first token?" — It is an attention sink: a learned place to put weight when nothing else is relevant, because softmax cannot output all zeros.
  • "What is an induction head?" — A pair of heads that together copy the token that followed an earlier occurrence of the current token; a key mechanism for in-context learning.
  • "How do you know a head matters?" — Ablate or patch it and measure the change in loss or behaviour; the weight map alone does not show importance.