Course Content
Transformer Architecture Q&A
6 sections · 60 lessons
Why is dot product (not cosine or Euclidean) used as similarity in attention?
What you need to know
The three options, as formulas
dot: s = q · kcosine: s = (q · k) / (|q| |k|) → always between -1 and 1Euclidean: s = -|q - k|² = 2 q·k - |q|² - |k|² → needs q·k anywayFor one query row, -|k|² differs per key and -|q|² is the same for all keys (softmax ignores it). So Euclidean attention is dot-product attention plus a penalty on large keys — more work for no clear gain.
Why cosine alone cannot focus
Softmax turns score gaps into probability ratios: a gap of Δ gives a ratio of exp(Δ). With cosine, the largest possible gap is 2.
1import numpy as np2n = 1000 # 1 matching key, 999 unrelated keys3cos = np.r_[1.0, np.zeros(n - 1)] # cosine scores live in [-1, 1]4w = np.exp(cos) / np.exp(cos).sum()5print(round(w[0], 4)) # 0.0027 — almost no focus67dot = np.r_[12.0, np.zeros(n - 1)] # dot product: magnitude can grow8w = np.exp(dot) / np.exp(dot).sum()9print(round(w[0], 4)) # 0.9939With 1,000 keys, a perfect cosine match gets 0.27% of the weight. Models that use cosine attention (Swin Transformer V2, for example) therefore multiply by a learned temperature — which reintroduces a magnitude.
Magnitude as a signal
With a dot product, a key vector that is long in the matching direction says "I strongly have this feature". The model can learn to make some keys large and others small. Cosine throws that away.
Controlling the downside
Dot products of random vectors grow with dimension: their variance is about d_k. Without care, scores become huge and the softmax saturates, giving near-zero gradients. Two fixes:
1/sqrt(d_k)scaling — from the original Transformer paper; keeps score variance near 1 at initialisation.- QK-norm — apply a normalisation with a learned scale to Q and K before the dot product (LayerNorm in ViT-22B, RMSNorm in most LLMs). It was used for stable training of ViT-22B (Dehghani et al., 2023) and appears in several 2025 open models, including OLMo 2, Qwen3 and Gemma 3. It is "cosine with a learned temperature", added where training stability needs it.
A real-life example
An e-commerce search system uses both ideas in different places. For retrieval, it embeds queries and products and uses cosine similarity, because it wants "is this product about the same thing?" regardless of how long or popular the product's text is — an embedding's length often tracks word frequency rather than relevance.
Inside the ranking Transformer, attention uses scaled dot products. There, a token in "red running shoes under 3000" must put most of its weight on one or two of dozens of tokens, and only the dot product can make that sharp. The team once tried cosine attention without a temperature in a small experiment; the attention maps came out nearly uniform, and ranking quality fell until they added a learned scale.
Follow-up questions to expect
- "Why divide by
sqrt(d_k)and notd_k?" — Each score is a sum ofd_kproducts with variance about 1, so its standard deviation grows assqrt(d_k); dividing by that keeps it near 1. - "What is QK-norm?" — Normalising Q and K per head before the dot product, with learned scales, which bounds the logits and prevents attention-logit blow-up in large training runs.
- "Is additive (Bahdanau) attention better?" — It uses a small MLP to score pairs and works well, but it cannot be done as one matrix multiply, so it is slower at scale.