Course Content
Transformer Architecture Q&A
6 sections · 60 lessons
What is a "context vector" in self-attention, and why is it the key output?
What you need to know
A small worked example
1import numpy as np2tokens = ["red", "running", "shoes"]3V = np.array([[1.0, 0.0, 0.0], # toy value vectors: [colour, sport, footwear]4 [0.0, 1.0, 0.0],5 [0.0, 0.1, 1.0]])6a = np.array([0.25, 0.45, 0.30]) # weights for the query at "shoes" (sum to 1)7c = a @ V # context vector for "shoes"8print(c.round(3)) # [0.25 0.48 0.3 ]The context vector for "shoes" now carries "sport" and "colour" information from its neighbours. After W_o and the residual add, the representation of "shoes" means "red running shoes", not just "shoes".
Why it is the key output
- It is the payload. Scores and softmax are the routing decision; the context vector is what gets routed.
- It is per token. Each query produces its own weights, so each token gets a summary tailored to what it is looking for. This is what separates attention from average pooling, where every token would get the same summary.
- It is the only cross-token path. The FFN, norms and residual adds all work on one position at a time. If a token needs to know anything about another token, it arrives through a context vector.
Where it goes next
c_i (per head, size d_head) → concat heads → W_o → add to residual stream → FFNIn a multi-head layer, each head computes its own context vector in its own subspace; they are concatenated and projected by W_o.
The older meaning of the term
In early sequence-to-sequence models (Sutskever et al., 2014), "the context vector" was one fixed-size vector summarising the whole source sentence — a bottleneck for long inputs. Bahdanau attention (2014) made a fresh context vector for every output step. Self-attention goes further: a context vector for every token, in every head, in every layer.
A real-life example
An e-commerce search-ranking model reads the query "red running shoes for men under 3000". A product whose title says "men's running shoe, red" should score high; a "red formal shoe" should not.
In the query encoder, the context vector for "shoes" draws heavily from "running", so its representation now leans towards sports footwear. The context vector for "3000" draws from "under", turning a bare number into "a price ceiling". When the team inspected a wrongly ranked query, they found the context vector of "shoes" drew mostly from "red" in early layers and the model ranked red formal shoes too high — a clue that the training data under-represented category words, which they fixed by adding more category-specific query pairs.
Follow-up questions to expect
- "Is the context vector the same as the attention weights?" — No. Weights are the mixing coefficients; the context vector is the mixed result, a vector of size
d_head. - "Why can a context vector be close to the token's own value?" — If the token attends mostly to itself,
a_iiis near 1, so it passes its own information through largely unchanged. - "Does the context vector replace the token's representation?" — No. It is added through the residual connection, so the original information is kept.