Course Content
Transformer Architecture Q&A
6 sections · 60 lessons
What roles do Query, Key, Value play in attention, and how do they interact mathematically?
What you need to know
The formulas and shapes
X (T, d_model) the token vectorsQ = X W_q (T, d_k)K = X W_k (T, d_k)V = X W_v (T, d_v)scores = Q Kᵀ / sqrt(d_k) (T, T) row i = query i vs every keyweights = softmax(scores, axis=-1) (T, T) each row sums to 1out = weights V (T, d_v) one new vector per tokenUsually d_k = d_v = d_model / h, where h is the number of heads.
1import numpy as np23rng = np.random.default_rng(0)4T, d_model, d_k, d_v = 3, 8, 4, 45X = rng.normal(size=(T, d_model)) # 3 tokens, 8 numbers each67W_q = rng.normal(size=(d_model, d_k)) / np.sqrt(d_model)8W_k = rng.normal(size=(d_model, d_k)) / np.sqrt(d_model)9W_v = rng.normal(size=(d_model, d_v)) / np.sqrt(d_model)1011Q, K, V = X @ W_q, X @ W_k, X @ W_v # each (3, 4)12scores = Q @ K.T / np.sqrt(d_k) # (3, 3): every query vs every key13weights = np.exp(scores - scores.max(axis=-1, keepdims=True))14weights /= weights.sum(axis=-1, keepdims=True)15out = weights @ V # (3, 4): one new vector per token1617print(scores.shape, weights.sum(axis=-1), out.shape)18# (3, 3) [1. 1. 1.] (3, 4)The d_k dimension disappears in Q Kᵀ — it is summed over — leaving a position-by-position table. Multiplying by V brings back a vector per token.
Why three separate projections
If you used X directly for all three, the score between tokens i and j would equal the score between j and i, and each token would match itself most strongly. Separate W_q and W_k let the model learn asymmetric relations: a pronoun can look for a noun without the noun looking for the pronoun. A separate W_v means what makes a token a good match can differ from what it passes on.
A real-life example
A speech-to-text system based on an encoder-decoder model such as Whisper processes 30 seconds of audio into 1,500 encoder vectors. When the decoder is about to write the next word, its query asks "which audio frames match the sound I need now?" The keys come from the 1,500 audio frames; the matching frames — say those around second 12 — get high weights. Their values carry the acoustic content that the decoder uses to write the word "Chennai".
Here the query and the keys come from different sequences (text and audio), so this is cross-attention, and the weight matrix for a 20-token transcript is 20 × 1,500, not square.
Follow-up questions to expect
- "Must
d_vequald_k?" — No.d_kmust match between Q and K for the dot product;d_vcan differ, though it is usually the same. - "Are Q, K, V computed separately in practice?" — Usually one fused matrix computes all three in one matmul, then the result is split.
- "Which of Q, K, V goes into the KV cache?" — K and V of past tokens, because future queries need them; old queries are never reused.