Course Content
Transformer Architecture Q&A
6 sections · 60 lessons
Why is softmax used to normalize attention scores instead of dividing by their sum?
What you need to know
The failure of plain normalisation
1import numpy as np2s = np.array([2.0, -3.0])3print(s / s.sum()) # [-2. 3.] — not an average4e = np.exp(s - s.max())5print((e / e.sum()).round(4)) # [0.9933 0.0067]6s = np.array([3.0, -2.0, -1.0])7print(s.sum()) # 0.0 — division by zeroDot-product scores are signed. Plain division can produce negative weights, weights above 1, or a division by zero. The output would no longer be a mix of the value vectors, and its size would jump around with the input.
The softmax formula
softmax(s)_j = exp(s_j - max(s)) / Σ_k exp(s_k - max(s))Subtracting max(s) changes nothing mathematically but stops exp from overflowing — the same trick FlashAttention's running max uses.
What softmax buys
- A valid distribution. Every weight is above 0, the row sums to 1, and the output is a convex combination of value vectors: its scale stays stable whatever the sequence length.
- Sharp focus. A score gap of
Δbecomes a weight ratio ofexp(Δ). A gap of 5 already means about 150× more weight. This is why scores are scaled by1/sqrt(d_k)— without it, gaps get so large that the softmax saturates and gradients vanish. - Competition. Raising one score lowers every other weight, so heads learn to choose.
- Clean masking.
exp(-inf) = 0, so masked keys get exactly zero and the rest renormalise. - Smooth gradients. It is differentiable everywhere, unlike a hard
argmax.
Alternatives in research
- Sigmoid attention replaces the row-wise softmax with an element-wise sigmoid (studied by Ramapuram et al., Apple, 2024); it needs careful normalisation to train well.
- Attention with a sink term adds a constant to the softmax denominator so a head can give low weight to every real token; gpt-oss uses a learned version.
- Linear attention drops the softmax to get linear-time cost, usually at some quality loss.
A real-life example
Think of splitting a festival bonus pool among five team members based on "scores". If some scores are negative, "divide by the total" can tell someone to pay money back, or blow up when the total is zero. Softmax is like first converting every score into a positive number of points, then splitting by points: everyone gets a share, the shares add up to the pool, and a clearly better score gets a clearly bigger share.
In production this shows up in e-commerce ranking too: listwise ranking losses apply softmax across a query's candidate products for the same reasons — scores from the model are signed, and the loss needs a proper distribution over candidates.
Follow-up questions to expect
- "Why subtract the max?" — To avoid overflow in
exp; it does not change the result because the same factor cancels in numerator and denominator. - "What happens if scores are very large?" — Softmax becomes nearly one-hot, gradients for the other positions vanish, and training stalls; that is what
1/sqrt(d_k)and QK-norm prevent. - "Could you use ReLU and then divide by the sum?" — If all scores are negative, the sum is zero; and ReLU gives no gradient to negative scores, so those keys never learn to become relevant.