Course Content
Transformer Architecture Q&A
6 sections · 60 lessons
Distinguish attention scores vs weights—when do scores become weights?
What you need to know
A worked row
Take one query and four keys, where key 4 is a padding token:
raw scores [ 2.0, 1.0, -1.0, 0.5]after mask [ 2.0, 1.0, -1.0, -inf] edit the SCORESsoftmax [0.705, 0.259, 0.035, 0.0] now they are WEIGHTS, sum = 1A score of -1.0 is not "negative attention". It only means a poor match, and it still gets a small positive weight (0.035).
What acts on which
| Operation | Acts on | Why there |
|---|---|---|
1 / sqrt(d_k) scaling | Scores | Controls softmax sharpness |
| Causal or padding mask | Scores (set to -inf) | Softmax then gives exactly 0 and renormalises the rest |
| ALiBi or relative position bias | Scores (added) | Shifts preferences before normalisation |
| Softmax | Scores to weights | The single conversion point |
| Attention dropout | Weights | Randomly drops mixing links in training |
Multiply by V | Weights | Weights are the mixing coefficients |
1import torch, torch.nn.functional as F2scores = torch.tensor([[2.0, 1.0, -1.0, 0.5]])3pad = torch.tensor([[False, False, False, True]]) # last token is padding4scores = scores.masked_fill(pad, float("-inf")) # edit SCORES5weights = scores.softmax(dim=-1) # scores become weights here6weights = F.dropout(weights, p=0.1, training=False) # edits WEIGHTS (training only)7print(weights, weights.sum().item())8# tensor([[0.7054, 0.2595, 0.0351, 0.0000]]) 1.0000001192092896The softmax runs along dim=-1, the key axis, so each query's row is a probability distribution over positions it may read from.
Scores in FlashAttention
Optimised kernels such as FlashAttention never store the full score or weight matrix. They compute scores tile by tile and keep a running maximum and running sum (an "online softmax"). The maths is identical, but you cannot read the weights back from a fused kernel — if you need attention maps for debugging, you must use a non-fused path.
A real-life example
A document classifier sorts support emails into ten categories using a BERT-style encoder. Emails are padded to 256 tokens in a batch. An engineer writing a custom layer zeroed the padding weights after softmax instead of masking scores before it.
The result: each real token's weights summed to less than 1, by an amount that depended on how much padding there was. The same email got a different prediction when batched with a long email than with short ones. Moving the mask to the scores (-inf before softmax) made predictions independent of batch composition. The weight rows sum to 1 again, and the bug disappears.
Follow-up questions to expect
- "When people plot attention, which do they plot?" — Weights, after softmax, usually one head in one layer.
- "Can a weight be exactly zero?" — Only for masked positions; softmax of a finite score is always positive.
- "Where does a temperature go?" — On the scores (divide before softmax), the same place as the
sqrt(d_k)scaling.