Course Content
Transformer Architecture Q&A
6 sections · 60 lessons
FFN as “memory bank + generalizer”—what does that mean in Transformers?
What you need to know
The idea comes from Geva et al., "Transformer Feed-Forward Layers Are Key-Value Memories" (2021). Rewrite the FFN:
h = act(W_up x) each row of W_up is a KEY: h_i is how strongly pattern i is presentout = W_down h each column of W_down is a VALUE: added in proportion to h_iA model with d_ff = 14,336 has 14,336 such memory slots per layer.
A toy memory, worked by hand
1import numpy as np23# 3 "memories": each key detects a pattern, each value is what gets written back4keys = np.array([[1.0, 0.0, 0.0], # fires on "capital of"5 [0.0, 1.0, 0.0], # fires on "currency of"6 [0.0, 0.0, 1.0]]) # fires on "language of"7values = np.array([[1.0, 0.0], # pushes towards city names8 [0.0, 1.0], # pushes towards currency names9 [0.5, 0.5]])1011x = np.array([0.9, 0.3, -0.4]) # mostly "capital of", a little "currency of"12h = np.maximum(keys @ x, 0) # activations: [0.9, 0.3, 0.0]13out = h @ values # weighted sum of values14print(h, out) # [0.9 0.3 0. ] [0.9 0.3]The input mostly matches key 1, so value 1 dominates the output. Key 3 does not fire (ReLU cuts its negative score to 0). Real layers have thousands of keys, in high dimensions, with patterns learned rather than labelled.
The two halves of the name
- Memory bank. Specific inputs trigger specific stored outputs. Studies found keys in early layers tend to match surface patterns (a word ending, an n-gram) and later ones more semantic patterns (a topic).
- Generaliser. The lookup is soft. A new sentence that partly matches several keys gets a blend of their values, not a failed lookup.
Evidence from model editing
ROME (Meng et al., 2022) located a fact such as "The Eiffel Tower is in Paris" mainly in middle-layer FFN weights, and changed it with a small, targeted weight update. MEMIT extended this to thousands of facts at once. This supports the view that the FFN stores factual associations, although edits can have side effects and do not always generalise to rephrased questions.
A real-life example
An Indian-language news app uses an LLM to write background paragraphs. After a state election, the model still names the previous Chief Minister. A developer who has read about ROME suggests editing the FFN weights to update the fact.
The senior engineer explains the trade-off using this lesson. Weight edits work for a single fact but can disturb nearby facts, must be redone for every model update, and are hard to audit. Facts that change — office-holders, prices, scores — belong in retrieval: fetch the current fact from a trusted database and put it in the prompt. The team uses RAG for current facts and relies on the FFN "memory" only for stable background knowledge.
Follow-up questions to expect
- "Is knowledge only in the FFN?" — No. Attention moves information to where it is needed, and facts are spread across layers; the FFN view is a useful approximation, not a full map.
- "Why do bigger models know more facts?" — More layers and a larger
d_ffmean more key-value slots. - "How does this relate to Mixture-of-Experts?" — MoE splits the FFN into several experts and routes each token to a few, adding memory capacity without adding per-token compute.