Course Content
Transformer Architecture Q&A
6 sections · 60 lessons
Walk through Bahdanau attention and why it was a turning point in sequence modeling?
What you need to know
The problem it solved
Before 2014, neural translation used an encoder RNN that read the source word by word and a decoder RNN that wrote the target. The only link between them was the encoder's final hidden state — one vector, say 1,000 numbers, for the whole sentence. Short sentences fit; long sentences lost detail, and translation quality fell as sentences got longer.
The mechanism
At decoder step t, with previous decoder state s_{t-1} and encoder states h_1 ... h_n:
e_tj = vᵀ tanh(W_s s_{t-1} + W_h h_j) score: how well h_j fits nowa_tj = softmax over j of e_tj weights sum to 1c_t = sum over j of a_tj * h_j context vector for this stepThe score is a tiny neural network (a tanh layer and a vector v), so it is called additive attention. Luong (2015) later used simpler multiplicative (dot-product) scores, and the Transformer's scaled dot-product is a descendant of that.
Worked by hand
Three source words, 2-dimensional encoder states:
h_1 = [1, 0] h_2 = [0, 1] h_3 = [1, 1]scores e = [2.0, 0.5, -1.0]weights a = softmax(e) = [0.786, 0.175, 0.039]context c = 0.786*[1,0] + 0.175*[0,1] + 0.039*[1,1] = [0.825, 0.214]At this step the decoder mostly "looks at" word 1. At the next step the scores change, and it looks elsewhere.
Why it was a turning point
- No fixed bottleneck. The decoder can reach any source word at every step.
- Learned alignment. Nobody labels which Hindi word matches which English word; the weights learn it through ordinary backpropagation.
- Interpretable. Plotting
a_tjshows a soft alignment grid. - A new primitive. Three years later, Attention Is All You Need (2017) kept attention, removed the RNN entirely, and made attention the core of the model.
A real-life example
An Indian-language news app in 2015 ran an RNN translation model for English-to-Hindi. Headlines translated well, but 40-word paragraphs lost names and numbers near the start: "The state government approved 1,250 new electric buses for Pune and Nagpur..." came back with the number missing.
Adding Bahdanau attention let the decoder look straight at "1,250" when it reached the point of writing the number, instead of hoping it survived inside one vector. Word order also improved, because Hindi puts the verb at the end: the decoder could attend back to the English verb near the start when it reached the end of the Hindi sentence.
Follow-up questions to expect
- "Additive versus dot-product attention?" — Additive uses a small network to score; dot-product just multiplies vectors, which is faster on GPUs and, with scaling, works as well.
- "What was still slow?" — The RNNs themselves process one step at a time, so training could not be parallelised across the sequence. The Transformer removed that limit.
- "Is Bahdanau attention self-attention?" — No, it is cross-attention: decoder state queries encoder states.