Transformer Architecture Q&A

Course Content

Transformer Architecture Q&A

6 sections · 60 lessons

Explain self-attention and how a token attends to others within a Transformer layer?


Where 'bank' looks in 'the river bank'0.140.630.23012theriverbankScores 0, 3, 1 divided by 2, then softmax.
The embedding for 'bank' is the same in every headline; putting 63% of its attention on 'river' is what pulls it toward the shore meaning.

What you need to know

Before attention, each token's vector only describes the token itself. The word "bank" has the same embedding whether it means a river bank or a money bank. Self-attention is how "bank" pulls in information from its neighbours so its vector reflects the sentence.

The three projections

Each token vector x is multiplied by three learned matrices:

  • query q = x W_q — what this token is looking for;
  • key k = x W_k — what this token offers to others;
  • value v = x W_v — the information it passes on if chosen.

One token, worked by hand

Sentence: "the river bank". Look only at the token bank. Assume d_k = 4, so the scale is sqrt(4) = 2.

Text
keys:   the   = [0, 0, 0, 1]        river = [1, 1, 0, 0]        bank  = [0, 1, 1, 0]query of bank = [2, 1, 0, 0]scores  = q . k  = [0, 3, 1]scaled  = / 2    = [0, 1.5, 0.5]weights = softmax = [0.14, 0.63, 0.23]values: the = [0, 0], river = [1, 0], bank = [0, 1]output  = 0.14*[0,0] + 0.63*[1,0] + 0.23*[0,1] = [0.63, 0.23]

"bank" gives 63% of its attention to "river", so its new vector now leans towards the river meaning. Every token does this at the same time; in matrix form it is softmax(Q Kᵀ / sqrt(d_k)) V.

Properties worth saying

  • Content-based. The weights depend on what tokens contain, not where they are. Position must be added separately.
  • All pairs. Every token scores every other token, so cost grows with T².
  • Direct paths. Token 1 and token 500 interact in one step. An RNN needed 499 steps, and information faded on the way.
  • Different every layer. Each layer has its own W_q, W_k, W_v, so early layers often attend to neighbours and later layers to more abstract links.

A real-life example

An Indian-language news app translates English headlines into Hindi. Two headlines contain "bank":

  • "Flood waters cross the river bank in Patna" should become किनारा (shore).
  • "Bank raises home loan rates" should become बैंक (the financial bank).

The embedding for "bank" is identical in both. In the first headline, attention gives "bank" a high weight on "river" and "flood"; in the second, on "loan" and "rates". After a few layers the two "bank" vectors are far apart, and the decoder picks different Hindi words. When the team inspected a wrong translation, the attention map showed "bank" attending mostly to itself because the context word was 30 tokens away in a long headline — a useful debugging signal, though not a full explanation.

Follow-up questions to expect

  • "What is the complexity?" — O(T² · d) time for the scores and O(T²) memory if the score matrix is stored; FlashAttention avoids storing it.
  • "How is it different from an RNN?" — An RNN passes a single state step by step; attention connects every pair directly and runs in parallel.
  • "Can a token attend to itself?" — Yes, and it often gives itself substantial weight.