Transformer Architecture Q&A

Course Content

Transformer Architecture Q&A

6 sections · 60 lessons

Compare "attention" vs "self-attention"—when are they the same and when do they differ?


Same formula, different source for QSelf-attention (encoder)• Q, K, V from the same 6 English tokens• Weight matrix is 6 x 6• Both directions allowed• Every layer of GPT and BERTCross-attention (decoder)• Q from the 8 Tamil tokens so far• K, V from the 6 encoder outputs• Weight matrix is 8 x 6, not square• Only source padding is masked
Only the origin of the queries changes, which is why a cross-attention map is rectangular and reads like a word alignment.

What you need to know

All three forms use the same maths: softmax(Q Kᵀ / sqrt(d_k)) V. The only question is where Q, K and V come from.

TypeQ fromK and V fromWeight matrix shapeWhere used
Self-attentionSequence XSame sequence XT × TEvery layer of GPT, BERT
Cross-attentionDecoder (target)Encoder output (source)T_target × T_sourceT5, Whisper, translation models
Classic RNN attentionDecoder RNN stateEncoder RNN states1 row per output stepBahdanau (2014), Luong (2015)

Self-attention

Each token looks at other tokens in its own sequence. In an encoder it sees both directions; in a decoder a causal mask limits it to earlier tokens.

Cross-attention

The decoder has a sequence of target tokens and the encoder has produced a sequence of source vectors. Cross-attention lets each target token ask "which source positions matter for me now?" The weight matrix is usually not square, because the two sequences have different lengths.

A decoder block in an encoder-decoder model has three sub-layers: masked self-attention, cross-attention, then the FFN.

A real-life example

An Indian-language news app translates English wire stories into Tamil with an encoder-decoder model (open Indic translation models such as AI4Bharat's IndicTrans2 use this design).

The English source "Heavy rain shuts schools in Chennai" is 6 tokens. The encoder runs self-attention over those 6 tokens with a 6 × 6 weight matrix, both directions allowed.

Suppose the Tamil output so far is 8 tokens. The decoder runs masked self-attention over its 8 tokens (8 × 8, lower-triangular) to keep its sentence fluent, then cross-attention from its 8 tokens to the 6 source tokens (8 × 6). When the decoder writes the Tamil word for Chennai, its cross-attention row puts most weight on source token "Chennai". Plotting that 8 × 6 matrix gives a rough word alignment, which the team uses to spot dropped words.

Follow-up questions to expect

  • "Does GPT have cross-attention?" — No. Decoder-only models only have causal self-attention; the "source" is simply earlier in the same sequence.
  • "Is cross-attention masked?" — Not causally. The whole source is known, so every target token may see every source token; only source padding is masked.
  • "Can K and V be cached in cross-attention?" — Yes, and they are computed once, because the encoder output does not change during decoding.