Course Content
Transformer Architecture Q&A
6 sections · 60 lessons
What roles do W_q, W_k, W_v play, and why must they be learned (not fixed)?
What you need to know
The three roles, with shapes
For one head, each is a d_model × d_head matrix (for example 768 × 64 in GPT-2 small):
q_i = x_i · W_q what token i is looking fork_j = x_j · W_k what token j advertisesv_j = x_j · W_v what token j contributes if chosenscore_ij = q_i · k_j / sqrt(d_head)They also pick a small subspace: each head sees only a 64-dimensional projection of the 768-dimensional stream, and can ignore everything else.
Why separate W_q and W_k: direction
1import torch2torch.manual_seed(0)3x = torch.randn(4, 16) # 4 tokens4W = torch.randn(16, 8)5same = (x @ W) @ (x @ W).T # W_q = W_k6print(torch.allclose(same, same.T)) # True: i→j equals j→i78W_q, W_k = torch.randn(16, 8), torch.randn(16, 8)9diff = (x @ W_q) @ (x @ W_k).T10print(torch.allclose(diff, diff.T)) # False: direction mattersA shared matrix makes score(i, j) = score(j, i). But an adjective should look at its noun more than the noun looks back; a closing bracket should find its opening bracket, not the reverse. Separate matrices allow this. Another way to see it: W_q · W_kᵀ is one learned d_model × d_model matrix of rank d_head that defines what "matching" means for that head.
Why a separate W_v: matching versus content
If W_v were the identity, whatever matched would be copied as-is. With a learned W_v, a head can match on one property and move a different one — find the subject of the sentence by its grammatical role, and copy its meaning. The pair W_v · W_o decides what information is moved and where it lands.
Why not fixed
With fixed or identity projections, the score becomes the raw similarity of embeddings — "tokens that look alike attend to each other". That cannot express "a verb looks for its subject", which are very different tokens. Learning the projections is what lets each head define its own notion of relevance.
A real-life example
A code-completion assistant reads:
1def shipping_fee(order_value, pincode):2 if order_value > 499:3 return 04 return base_rate(pincode) + ...At the second pincode, one head's query asks "where was this name defined?". The key of the first pincode (the parameter) advertises "I am a parameter name in this function's signature", which matches. The value it passes is different again: information like "this is a location code, likely a string", which helps the model suggest a sensible next call.
A fixed similarity would match pincode with every other occurrence of pincode equally, including an unrelated one in another function earlier in the file. Learned projections let the key encode scope as well as spelling.
Follow-up questions to expect
- "Why are they
d_model × d_headand notd_model × d_model?" — Each head works in a low-dimensional subspace; stacked across all heads they form oned_model × d_modelmatrix. - "Can
W_kandW_vbe shared across heads?" — Yes, that is MQA and GQA, which share K and V across groups of query heads to shrink the KV cache. - "Do they have biases?" — GPT-2 has them; Llama-style models do not. Some Qwen models kept a bias on Q/K/V only.