Course Content
LLMs Deep Dive
10 sections · 40 lessons
What is the softmax function and why is it used in attention?
What you need to know
softmax(z_i) = exp(z_i) / sum over j of exp(z_j)A worked example
Scores [2.0, 1.0, 0.1]:
exp: 7.39, 2.72, 1.11 sum = 11.21softmax: 0.66, 0.24, 0.10 sum = 1.00A score gap of 1.0 becomes a ratio of about 2.7 (e to the power 1). Gaps become ratios, which is what "sharpening" means.
Why attention needs it
- Positive weights — a negative mixing weight would subtract information, which makes no sense for "how much to look at this token".
- Weights sum to 1 — the output is an average, so its size does not grow with the number of tokens. Without this, a 10,000-token input would produce much larger outputs than a 10-token input.
- Differentiable — unlike picking the single best token (argmax), softmax has useful gradients everywhere, so the model can learn where to attend.
Where softmax appears in an LLM
- Inside every attention head, once per query row.
- At the output, over the full vocabulary (e.g. 128,000 logits) to give next-token probabilities.
- With temperature, the logits are divided by T before softmax: T below 1 sharpens the distribution, T above 1 flattens it.
Numerical stability
exp(1000) overflows a float. Since softmax gives the same result if you subtract a constant from every score, implementations subtract the maximum first:
1import numpy as np23def softmax(z):4 z = z - z.max() # largest becomes 0, no overflow5 e = np.exp(z)6 return e / e.sum()78print(softmax(np.array([1000.0, 999.0, 998.0])).round(3)) # [0.665 0.245 0.09]Saturation
When one score is much larger than the rest (e.g. 30 versus 5), softmax outputs almost exactly 1 and 0. Its gradient becomes nearly zero, so the model stops learning which token to prefer. This is why attention scores are scaled by sqrt(d_k).
A real-life example
An e-commerce assistant handles a code-mixed query and is about to write the next word after "Aapka order kal". The output layer gives logits for three candidates: " deliver" 3.2, " ship" 2.1, " cancel" 0.4. Softmax (restricted to these three for simplicity) gives about 0.72, 0.24 and 0.04.
With sampling at temperature 1, "cancel" would appear about 4 times in 100 messages — telling a customer their order will be cancelled when it will not. The team sets a low temperature for this templated flow, which sharpens the distribution so " deliver" is chosen almost every time, and they generate the delivery date from the order system rather than from the model.
Follow-up questions to expect
- "Why exp and not just divide by the sum?" — Raw scores can be negative, and plain division would not sharpen differences; exp makes everything positive and turns gaps into ratios.
- "What is log-softmax?" — The log of softmax, computed directly for stability; it is what cross-entropy loss uses.
- "Are there alternatives to softmax in attention?" — Research has tried sigmoid and linear attention to avoid the normalisation cost, but standard softmax remains the default in frontier LLMs.