Course Content
LLMs Deep Dive
10 sections · 40 lessons
What is attention in Transformer models?
What you need to know
The idea in plain words
When you read "it", you look back to find what "it" refers to. Attention is a learned, soft version of that: instead of choosing one earlier word, each token takes a mix of all allowed tokens, weighted by relevance.
The formula
Attention(Q, K, V) = softmax(Q · K^T / sqrt(d_k)) · VQ · K^T— a score for every query–key pair (dot products)./ sqrt(d_k)— scaling so scores do not get too large (d_k is the key size, e.g. 128).softmax— turns each row of scores into weights between 0 and 1 that sum to 1.· V— each token's output is the weighted average of the value vectors.
Three kinds of attention
- Self-attention — queries, keys and values all come from the same sequence. Used in every LLM layer.
- Cross-attention — queries from one sequence, keys and values from another (a decoder reading an encoder's output in translation, or text reading image features).
- Causal (masked) attention — self-attention where future positions are blocked, so a decoder cannot see the answer it is trying to predict.
Why it works so well
Attention is content-based: which tokens are relevant is decided by what they contain, not by a fixed window. The same layer can link a pronoun to a noun 3 words back or a clause reference to a definition 3,000 tokens back. Stacking dozens of layers lets the model build up meaning step by step: early layers often track nearby words, later ones track meaning and long-range structure.
A real-life example
A bank's support bot receives a code-mixed message: "Maine kal payment kiya but it failed, paisa kat gaya" ("I made a payment yesterday but it failed, money got deducted"). To answer correctly, the model must connect "it" to "payment" and "paisa kat gaya" (money deducted) to the same transaction.
In one attention head, the query for "it" might give weights like this (illustrative):
Maine 0.03 | kal 0.05 | payment 0.71 | kiya 0.06 | but 0.04 | it 0.11So 71% of the output vector for "it" comes from "payment". Later layers use that enriched vector to recognise "failed payment with debit", and the bot opens the failed-transaction flow instead of a generic reply. The same mechanism works across languages because it matches meaning in vectors, not exact words.
Follow-up questions to expect
- "Why three separate vectors instead of comparing embeddings directly?" — Separate learned projections let a token look for one thing (query) while advertising something else (key) and passing on a third (value). That is far more flexible.
- "Are attention weights an explanation of the model's decision?" — Only partly. They show where one head looked, but many heads and layers combine, so they are not a reliable explanation on their own.
- "What is the cost of attention?" — O(n squared) in sequence length for compute; with FlashAttention, memory is linear.