Transformer Architecture Q&A

Course Content

Transformer Architecture Q&A

6 sections · 60 lessons

Walk through Bahdanau attention and why it was a turning point in sequence modeling?


What you need to know

The problem it solved

Before 2014, neural translation used an encoder RNN that read the source word by word and a decoder RNN that wrote the target. The only link between them was the encoder's final hidden state — one vector, say 1,000 numbers, for the whole sentence. Short sentences fit; long sentences lost detail, and translation quality fell as sentences got longer.

The mechanism

At decoder step t, with previous decoder state s_{t-1} and encoder states h_1 ... h_n:

Text
e_tj = vᵀ tanh(W_s s_{t-1} + W_h h_j)    score: how well h_j fits nowa_tj = softmax over j of e_tj              weights sum to 1c_t  = sum over j of a_tj * h_j            context vector for this step

The score is a tiny neural network (a tanh layer and a vector v), so it is called additive attention. Luong (2015) later used simpler multiplicative (dot-product) scores, and the Transformer's scaled dot-product is a descendant of that.

Worked by hand

Three source words, 2-dimensional encoder states:

Text
h_1 = [1, 0]   h_2 = [0, 1]   h_3 = [1, 1]scores e = [2.0, 0.5, -1.0]weights a = softmax(e) = [0.786, 0.175, 0.039]context c = 0.786*[1,0] + 0.175*[0,1] + 0.039*[1,1] = [0.825, 0.214]

At this step the decoder mostly "looks at" word 1. At the next step the scores change, and it looks elsewhere.

Why it was a turning point

  • No fixed bottleneck. The decoder can reach any source word at every step.
  • Learned alignment. Nobody labels which Hindi word matches which English word; the weights learn it through ordinary backpropagation.
  • Interpretable. Plotting a_tj shows a soft alignment grid.
  • A new primitive. Three years later, Attention Is All You Need (2017) kept attention, removed the RNN entirely, and made attention the core of the model.

A real-life example

An Indian-language news app in 2015 ran an RNN translation model for English-to-Hindi. Headlines translated well, but 40-word paragraphs lost names and numbers near the start: "The state government approved 1,250 new electric buses for Pune and Nagpur..." came back with the number missing.

Adding Bahdanau attention let the decoder look straight at "1,250" when it reached the point of writing the number, instead of hoping it survived inside one vector. Word order also improved, because Hindi puts the verb at the end: the decoder could attend back to the English verb near the start when it reached the end of the Hindi sentence.

Follow-up questions to expect

  • "Additive versus dot-product attention?" — Additive uses a small network to score; dot-product just multiplies vectors, which is faster on GPUs and, with scaling, works as well.
  • "What was still slow?" — The RNNs themselves process one step at a time, so training could not be parallelised across the sequence. The Transformer removed that limit.
  • "Is Bahdanau attention self-attention?" — No, it is cross-attention: decoder state queries encoder states.