Transformer Architecture Q&A

Course Content

Transformer Architecture Q&A

6 sections · 60 lessons

Why is dot product (not cosine or Euclidean) used as similarity in attention?


What you need to know

The three options, as formulas

Text
dot:        s = q · kcosine:     s = (q · k) / (|q| |k|)                 → always between -1 and 1Euclidean:  s = -|q - k|² = 2 q·k - |q|² - |k|²     → needs q·k anyway

For one query row, -|k|² differs per key and -|q|² is the same for all keys (softmax ignores it). So Euclidean attention is dot-product attention plus a penalty on large keys — more work for no clear gain.

Why cosine alone cannot focus

Softmax turns score gaps into probability ratios: a gap of Δ gives a ratio of exp(Δ). With cosine, the largest possible gap is 2.

Python
import numpy as npn = 1000                                   # 1 matching key, 999 unrelated keyscos = np.r_[1.0, np.zeros(n - 1)]          # cosine scores live in [-1, 1]w = np.exp(cos) / np.exp(cos).sum()print(round(w[0], 4))                      # 0.0027 — almost no focusdot = np.r_[12.0, np.zeros(n - 1)]         # dot product: magnitude can groww = np.exp(dot) / np.exp(dot).sum()print(round(w[0], 4))                      # 0.9939

With 1,000 keys, a perfect cosine match gets 0.27% of the weight. Models that use cosine attention (Swin Transformer V2, for example) therefore multiply by a learned temperature — which reintroduces a magnitude.

Magnitude as a signal

With a dot product, a key vector that is long in the matching direction says "I strongly have this feature". The model can learn to make some keys large and others small. Cosine throws that away.

Controlling the downside

Dot products of random vectors grow with dimension: their variance is about d_k. Without care, scores become huge and the softmax saturates, giving near-zero gradients. Two fixes:

  • 1/sqrt(d_k) scaling — from the original Transformer paper; keeps score variance near 1 at initialisation.
  • QK-norm — apply a normalisation with a learned scale to Q and K before the dot product (LayerNorm in ViT-22B, RMSNorm in most LLMs). It was used for stable training of ViT-22B (Dehghani et al., 2023) and appears in several 2025 open models, including OLMo 2, Qwen3 and Gemma 3. It is "cosine with a learned temperature", added where training stability needs it.

A real-life example

An e-commerce search system uses both ideas in different places. For retrieval, it embeds queries and products and uses cosine similarity, because it wants "is this product about the same thing?" regardless of how long or popular the product's text is — an embedding's length often tracks word frequency rather than relevance.

Inside the ranking Transformer, attention uses scaled dot products. There, a token in "red running shoes under 3000" must put most of its weight on one or two of dozens of tokens, and only the dot product can make that sharp. The team once tried cosine attention without a temperature in a small experiment; the attention maps came out nearly uniform, and ranking quality fell until they added a learned scale.

Follow-up questions to expect

  • "Why divide by sqrt(d_k) and not d_k?" — Each score is a sum of d_k products with variance about 1, so its standard deviation grows as sqrt(d_k); dividing by that keeps it near 1.
  • "What is QK-norm?" — Normalising Q and K per head before the dot product, with learned scales, which bounds the logits and prevents attention-logit blow-up in large training runs.
  • "Is additive (Bahdanau) attention better?" — It uses a small MLP to score pairs and works well, but it cannot be done as one matrix multiply, so it is slower at scale.