Transformer Architecture Q&A

Course Content

Transformer Architecture Q&A

6 sections · 60 lessons

Difference between model "weight parameters" and "attention weights" at inference?


What you need to know

Weight parametersAttention weights
What they arelearned tensorssoftmax(QKᵀ / sqrt(d)) for this input
Depend on the inputnoyes
Lifetimepermanent (the checkpoint)one forward pass
Counted in "8B"yesno
Changed bytraining, fine-tuning, LoRA, quantisationa different prompt

Put numbers on it

Llama-3-8B has about 8.0 billion weight parameters. For one 4,096-token prompt, the attention weights across all layers and heads would be:

Text
32 layers × 32 heads × 4,096 × 4,096 ≈ 17.2 billion numbers

That is more numbers than the model has parameters, created for one request and then discarded. FlashAttention never writes the full matrix at all. That is why you should never try to "save the attention weights" for a production chatbot: they are enormous and specific to one input.

What each side is touched by

  • Parameter side: training, fine-tuning, LoRA adapters, pruning, quantisation (int8, int4, FP8), weight tying.
  • Activation side: attention visualisation, the KV cache (which stores K and V activations, not attention weights), FlashAttention, masking.

The KV cache is a common source of confusion. It is activation data like attention weights — per request, input-dependent — but it stores keys and values, not the softmax matrix.

Getting attention weights out

In Hugging Face transformers, pass output_attentions=True. Fused implementations such as SDPA and FlashAttention do not produce the matrix, so the model must use the "eager" implementation, which is slower and uses much more memory at long context.

A real-life example

A product manager on a chatbot team serving a million users asks: "Since the model adjusts its attention weights for each conversation, is it learning from our users?" The engineer's answer: no. The attention weights are recomputed from scratch for each request and discarded; the 8B parameters on the GPU are identical before and after every conversation, and are shared read-only by all users.

What does carry information between turns is the conversation text the app sends back in (and the KV cache for that one session). If the team wants the model itself to change, they must fine-tune — train new parameters offline, evaluate them and deploy a new checkpoint. This distinction also matters for privacy reviews: user data lives in logs and caches, not in the weights, unless you train on it.

Follow-up questions to expect

  • "Is the KV cache part of the weights?" — No. It holds per-request K and V activations and is discarded when the session ends.
  • "Does in-context learning change the weights?" — No. The model adapts its behaviour through attention over the prompt; the parameters are unchanged.
  • "What does LoRA change?" — It adds small trained low-rank matrices to some weight parameters, such as W_q and W_v; attention weights are still computed fresh per input.