Course Content
Transformer Architecture Q&A
6 sections · 60 lessons
Describe the feed-forward network in a Transformer block—what it computes and why it exists?
What you need to know
FFN(x) = W_down · act(W_up · x + b_up) + b_downd_model -> d_ff -> d_model d_ff = 4 × d_model in the original"Position-wise" means the same weights are applied to each token vector on its own, like a 1×1 convolution over the sequence.
1import torch2from torch import nn34class FFN(nn.Module):5 def __init__(self, d_model=768, d_ff=3072):6 super().__init__()7 self.up = nn.Linear(d_model, d_ff) # expand 768 -> 30728 self.down = nn.Linear(d_ff, d_model) # project back 3072 -> 768910 def forward(self, x): # x: (B, T, d_model)11 return self.down(nn.functional.gelu(self.up(x)))1213ffn = FFN()14x = torch.randn(2, 5, 768)15print(ffn(x).shape, sum(p.numel() for p in ffn.parameters()))16# torch.Size([2, 5, 768]) 472243217# each position is processed alone: changing token 0 leaves token 1's output unchanged18y1 = ffn(x); x2 = x.clone(); x2[:, 0] += 1.019print(torch.allclose(ffn(x2)[:, 1:], y1[:, 1:]))20# TrueThe last check proves the point: changing token 0 does not change any other token's FFN output.
Why it is needed
Attention output is a weighted average of value vectors — a mostly linear operation. Stack only averages and the model stays close to linear. The FFN supplies the non-linearity and the capacity to transform the gathered information. A common summary: attention decides what to gather; the FFN decides what to do with it.
Where the parameters are
For GPT-2 small (d_model = 768): attention has 4 × 768² ≈ 2.36M weights per layer, the FFN 8 × 768² ≈ 4.72M. The FFN is two-thirds of the block.
Modern models shift this further. In Llama 3 8B, d_model = 4096, the SwiGLU FFN has hidden size 14,336, and GQA shrinks the K and V projections:
FFN 3 × 4096 × 14336 ≈ 176M per layerattention 4096² (Q) + 4096² (O) + 2 × 4096 × 1024 (K, V) ≈ 42M per layerFFN share ≈ 81%So in a 2026 LLM, the FFN is the dominant cost in parameters and in FLOPs at short context. It is also what Mixture-of-Experts models replicate: each "expert" is an FFN.
A real-life example
A company wants to host its own code-completion model on a single GPU and must cut the model's size. The engineer's first idea is to reduce attention heads. Using the numbers above, she sees that attention is under a fifth of the per-layer weights; halving it barely moves memory. Reducing d_ff from 14,336 to 11,008 saves far more.
They later move to a Mixture-of-Experts version: 8 FFN experts per layer, 2 active per token. Total parameters grow, but each token runs through only 2 FFNs, so per-token compute stays close to the dense model. Both decisions depend on knowing that the FFN is where the weights live.
Follow-up questions to expect
- "Why expand to 4× and come back?" — The wider hidden layer gives many more non-linear feature detectors; projecting back keeps the residual stream width fixed.
- "Does the FFN see other tokens?" — No. Any context it uses was already mixed into the vector by attention.
- "What replaced ReLU in the FFN?" — GELU in BERT and GPT-2; SwiGLU in most LLMs since 2023.