Course Content
Transformer Architecture Q&A
6 sections · 60 lessons
Compare "attention" vs "self-attention"—when are they the same and when do they differ?
What you need to know
All three forms use the same maths: softmax(Q Kᵀ / sqrt(d_k)) V. The only question is where Q, K and V come from.
| Type | Q from | K and V from | Weight matrix shape | Where used |
|---|---|---|---|---|
| Self-attention | Sequence X | Same sequence X | T × T | Every layer of GPT, BERT |
| Cross-attention | Decoder (target) | Encoder output (source) | T_target × T_source | T5, Whisper, translation models |
| Classic RNN attention | Decoder RNN state | Encoder RNN states | 1 row per output step | Bahdanau (2014), Luong (2015) |
Self-attention
Each token looks at other tokens in its own sequence. In an encoder it sees both directions; in a decoder a causal mask limits it to earlier tokens.
Cross-attention
The decoder has a sequence of target tokens and the encoder has produced a sequence of source vectors. Cross-attention lets each target token ask "which source positions matter for me now?" The weight matrix is usually not square, because the two sequences have different lengths.
A decoder block in an encoder-decoder model has three sub-layers: masked self-attention, cross-attention, then the FFN.
A real-life example
An Indian-language news app translates English wire stories into Tamil with an encoder-decoder model (open Indic translation models such as AI4Bharat's IndicTrans2 use this design).
The English source "Heavy rain shuts schools in Chennai" is 6 tokens. The encoder runs self-attention over those 6 tokens with a 6 × 6 weight matrix, both directions allowed.
Suppose the Tamil output so far is 8 tokens. The decoder runs masked self-attention over its 8 tokens (8 × 8, lower-triangular) to keep its sentence fluent, then cross-attention from its 8 tokens to the 6 source tokens (8 × 6). When the decoder writes the Tamil word for Chennai, its cross-attention row puts most weight on source token "Chennai". Plotting that 8 × 6 matrix gives a rough word alignment, which the team uses to spot dropped words.
Follow-up questions to expect
- "Does GPT have cross-attention?" — No. Decoder-only models only have causal self-attention; the "source" is simply earlier in the same sequence.
- "Is cross-attention masked?" — Not causally. The whole source is known, so every target token may see every source token; only source padding is masked.
- "Can K and V be cached in cross-attention?" — Yes, and they are computed once, because the encoder output does not change during decoding.