LLMs Deep Dive

Course Content

LLMs Deep Dive

10 sections · 40 lessons

What is the softmax function and why is it used in attention?


What you need to know

Text
softmax(z_i) = exp(z_i) / sum over j of exp(z_j)

A worked example

Scores [2.0, 1.0, 0.1]:

Text
exp:      7.39, 2.72, 1.11      sum = 11.21softmax:  0.66, 0.24, 0.10      sum = 1.00

A score gap of 1.0 becomes a ratio of about 2.7 (e to the power 1). Gaps become ratios, which is what "sharpening" means.

Why attention needs it

  • Positive weights — a negative mixing weight would subtract information, which makes no sense for "how much to look at this token".
  • Weights sum to 1 — the output is an average, so its size does not grow with the number of tokens. Without this, a 10,000-token input would produce much larger outputs than a 10-token input.
  • Differentiable — unlike picking the single best token (argmax), softmax has useful gradients everywhere, so the model can learn where to attend.

Where softmax appears in an LLM

  1. Inside every attention head, once per query row.
  2. At the output, over the full vocabulary (e.g. 128,000 logits) to give next-token probabilities.
  3. With temperature, the logits are divided by T before softmax: T below 1 sharpens the distribution, T above 1 flattens it.

Numerical stability

exp(1000) overflows a float. Since softmax gives the same result if you subtract a constant from every score, implementations subtract the maximum first:

Python
import numpy as npdef softmax(z):    z = z - z.max()          # largest becomes 0, no overflow    e = np.exp(z)    return e / e.sum()print(softmax(np.array([1000.0, 999.0, 998.0])).round(3))  # [0.665 0.245 0.09]

Saturation

When one score is much larger than the rest (e.g. 30 versus 5), softmax outputs almost exactly 1 and 0. Its gradient becomes nearly zero, so the model stops learning which token to prefer. This is why attention scores are scaled by sqrt(d_k).

A real-life example

An e-commerce assistant handles a code-mixed query and is about to write the next word after "Aapka order kal". The output layer gives logits for three candidates: " deliver" 3.2, " ship" 2.1, " cancel" 0.4. Softmax (restricted to these three for simplicity) gives about 0.72, 0.24 and 0.04.

With sampling at temperature 1, "cancel" would appear about 4 times in 100 messages — telling a customer their order will be cancelled when it will not. The team sets a low temperature for this templated flow, which sharpens the distribution so " deliver" is chosen almost every time, and they generate the delivery date from the order system rather than from the model.

Follow-up questions to expect

  • "Why exp and not just divide by the sum?" — Raw scores can be negative, and plain division would not sharpen differences; exp makes everything positive and turns gaps into ratios.
  • "What is log-softmax?" — The log of softmax, computed directly for stability; it is what cross-entropy loss uses.
  • "Are there alternatives to softmax in attention?" — Research has tried sigmoid and linear attention to avoid the normalisation cost, but standard softmax remains the default in frontier LLMs.