Course Content
Transformer Architecture Q&A
6 sections · 60 lessons
Difference between model "weight parameters" and "attention weights" at inference?
What you need to know
| Weight parameters | Attention weights | |
|---|---|---|
| What they are | learned tensors | softmax(QKᵀ / sqrt(d)) for this input |
| Depend on the input | no | yes |
| Lifetime | permanent (the checkpoint) | one forward pass |
| Counted in "8B" | yes | no |
| Changed by | training, fine-tuning, LoRA, quantisation | a different prompt |
Put numbers on it
Llama-3-8B has about 8.0 billion weight parameters. For one 4,096-token prompt, the attention weights across all layers and heads would be:
32 layers × 32 heads × 4,096 × 4,096 ≈ 17.2 billion numbersThat is more numbers than the model has parameters, created for one request and then discarded. FlashAttention never writes the full matrix at all. That is why you should never try to "save the attention weights" for a production chatbot: they are enormous and specific to one input.
What each side is touched by
- Parameter side: training, fine-tuning, LoRA adapters, pruning, quantisation (int8, int4, FP8), weight tying.
- Activation side: attention visualisation, the KV cache (which stores K and V activations, not attention weights), FlashAttention, masking.
The KV cache is a common source of confusion. It is activation data like attention weights — per request, input-dependent — but it stores keys and values, not the softmax matrix.
Getting attention weights out
In Hugging Face transformers, pass output_attentions=True. Fused implementations such as SDPA and FlashAttention do not produce the matrix, so the model must use the "eager" implementation, which is slower and uses much more memory at long context.
A real-life example
A product manager on a chatbot team serving a million users asks: "Since the model adjusts its attention weights for each conversation, is it learning from our users?" The engineer's answer: no. The attention weights are recomputed from scratch for each request and discarded; the 8B parameters on the GPU are identical before and after every conversation, and are shared read-only by all users.
What does carry information between turns is the conversation text the app sends back in (and the KV cache for that one session). If the team wants the model itself to change, they must fine-tune — train new parameters offline, evaluate them and deploy a new checkpoint. This distinction also matters for privacy reviews: user data lives in logs and caches, not in the weights, unless you train on it.
Follow-up questions to expect
- "Is the KV cache part of the weights?" — No. It holds per-request K and V activations and is discarded when the session ends.
- "Does in-context learning change the weights?" — No. The model adapts its behaviour through attention over the prompt; the parameters are unchanged.
- "What does LoRA change?" — It adds small trained low-rank matrices to some weight parameters, such as
W_qandW_v; attention weights are still computed fresh per input.