Course Content
Transformer Architecture Q&A
6 sections · 60 lessons
Explain self-attention and how a token attends to others within a Transformer layer?
What you need to know
Before attention, each token's vector only describes the token itself. The word "bank" has the same embedding whether it means a river bank or a money bank. Self-attention is how "bank" pulls in information from its neighbours so its vector reflects the sentence.
The three projections
Each token vector x is multiplied by three learned matrices:
- query
q = x W_q— what this token is looking for; - key
k = x W_k— what this token offers to others; - value
v = x W_v— the information it passes on if chosen.
One token, worked by hand
Sentence: "the river bank". Look only at the token bank. Assume d_k = 4, so the scale is sqrt(4) = 2.
keys: the = [0, 0, 0, 1] river = [1, 1, 0, 0] bank = [0, 1, 1, 0]query of bank = [2, 1, 0, 0]scores = q . k = [0, 3, 1]scaled = / 2 = [0, 1.5, 0.5]weights = softmax = [0.14, 0.63, 0.23]values: the = [0, 0], river = [1, 0], bank = [0, 1]output = 0.14*[0,0] + 0.63*[1,0] + 0.23*[0,1] = [0.63, 0.23]"bank" gives 63% of its attention to "river", so its new vector now leans towards the river meaning. Every token does this at the same time; in matrix form it is softmax(Q Kᵀ / sqrt(d_k)) V.
Properties worth saying
- Content-based. The weights depend on what tokens contain, not where they are. Position must be added separately.
- All pairs. Every token scores every other token, so cost grows with
T². - Direct paths. Token 1 and token 500 interact in one step. An RNN needed 499 steps, and information faded on the way.
- Different every layer. Each layer has its own
W_q,W_k,W_v, so early layers often attend to neighbours and later layers to more abstract links.
A real-life example
An Indian-language news app translates English headlines into Hindi. Two headlines contain "bank":
- "Flood waters cross the river bank in Patna" should become किनारा (shore).
- "Bank raises home loan rates" should become बैंक (the financial bank).
The embedding for "bank" is identical in both. In the first headline, attention gives "bank" a high weight on "river" and "flood"; in the second, on "loan" and "rates". After a few layers the two "bank" vectors are far apart, and the decoder picks different Hindi words. When the team inspected a wrong translation, the attention map showed "bank" attending mostly to itself because the context word was 30 tokens away in a long headline — a useful debugging signal, though not a full explanation.
Follow-up questions to expect
- "What is the complexity?" —
O(T² · d)time for the scores andO(T²)memory if the score matrix is stored; FlashAttention avoids storing it. - "How is it different from an RNN?" — An RNN passes a single state step by step; attention connects every pair directly and runs in parallel.
- "Can a token attend to itself?" — Yes, and it often gives itself substantial weight.