Transformer Architecture Q&A

Course Content

Transformer Architecture Q&A

6 sections · 60 lessons

What is a "context vector" in self-attention, and why is it the key output?


Weights from the query at 'shoes'0.250.450.30012redrunningshoesContext vector = 0.25 red + 0.45 running + 0.30 shoes, mixing their value vectors.
After this mix, 'shoes' carries 'running' and 'red' with it, which is how the ranker stops matching formal shoes.

What you need to know

A small worked example

Python
import numpy as nptokens = ["red", "running", "shoes"]V = np.array([[1.0, 0.0, 0.0],     # toy value vectors: [colour, sport, footwear]              [0.0, 1.0, 0.0],              [0.0, 0.1, 1.0]])a = np.array([0.25, 0.45, 0.30])   # weights for the query at "shoes" (sum to 1)c = a @ V                          # context vector for "shoes"print(c.round(3))                  # [0.25  0.48  0.3 ]

The context vector for "shoes" now carries "sport" and "colour" information from its neighbours. After W_o and the residual add, the representation of "shoes" means "red running shoes", not just "shoes".

Why it is the key output

  • It is the payload. Scores and softmax are the routing decision; the context vector is what gets routed.
  • It is per token. Each query produces its own weights, so each token gets a summary tailored to what it is looking for. This is what separates attention from average pooling, where every token would get the same summary.
  • It is the only cross-token path. The FFN, norms and residual adds all work on one position at a time. If a token needs to know anything about another token, it arrives through a context vector.

Where it goes next

Text
c_i (per head, size d_head) → concat heads → W_o → add to residual stream → FFN

In a multi-head layer, each head computes its own context vector in its own subspace; they are concatenated and projected by W_o.

The older meaning of the term

In early sequence-to-sequence models (Sutskever et al., 2014), "the context vector" was one fixed-size vector summarising the whole source sentence — a bottleneck for long inputs. Bahdanau attention (2014) made a fresh context vector for every output step. Self-attention goes further: a context vector for every token, in every head, in every layer.

A real-life example

An e-commerce search-ranking model reads the query "red running shoes for men under 3000". A product whose title says "men's running shoe, red" should score high; a "red formal shoe" should not.

In the query encoder, the context vector for "shoes" draws heavily from "running", so its representation now leans towards sports footwear. The context vector for "3000" draws from "under", turning a bare number into "a price ceiling". When the team inspected a wrongly ranked query, they found the context vector of "shoes" drew mostly from "red" in early layers and the model ranked red formal shoes too high — a clue that the training data under-represented category words, which they fixed by adding more category-specific query pairs.

Follow-up questions to expect

  • "Is the context vector the same as the attention weights?" — No. Weights are the mixing coefficients; the context vector is the mixed result, a vector of size d_head.
  • "Why can a context vector be close to the token's own value?" — If the token attends mostly to itself, a_ii is near 1, so it passes its own information through largely unchanged.
  • "Does the context vector replace the token's representation?" — No. It is added through the residual connection, so the original information is kept.