Transformer Architecture Q&A

Course Content

Transformer Architecture Q&A

6 sections · 60 lessons

Describe the feed-forward network in a Transformer block—what it computes and why it exists?


What you need to know

Text
FFN(x) = W_down · act(W_up · x + b_up) + b_downd_model -> d_ff -> d_model          d_ff = 4 × d_model in the original

"Position-wise" means the same weights are applied to each token vector on its own, like a 1×1 convolution over the sequence.

Python
import torchfrom torch import nnclass FFN(nn.Module):    def __init__(self, d_model=768, d_ff=3072):        super().__init__()        self.up = nn.Linear(d_model, d_ff)      # expand 768 -> 3072        self.down = nn.Linear(d_ff, d_model)    # project back 3072 -> 768    def forward(self, x):                       # x: (B, T, d_model)        return self.down(nn.functional.gelu(self.up(x)))ffn = FFN()x = torch.randn(2, 5, 768)print(ffn(x).shape, sum(p.numel() for p in ffn.parameters()))# torch.Size([2, 5, 768]) 4722432# each position is processed alone: changing token 0 leaves token 1's output unchangedy1 = ffn(x); x2 = x.clone(); x2[:, 0] += 1.0print(torch.allclose(ffn(x2)[:, 1:], y1[:, 1:]))# True

The last check proves the point: changing token 0 does not change any other token's FFN output.

Why it is needed

Attention output is a weighted average of value vectors — a mostly linear operation. Stack only averages and the model stays close to linear. The FFN supplies the non-linearity and the capacity to transform the gathered information. A common summary: attention decides what to gather; the FFN decides what to do with it.

Where the parameters are

For GPT-2 small (d_model = 768): attention has 4 × 768² ≈ 2.36M weights per layer, the FFN 8 × 768² ≈ 4.72M. The FFN is two-thirds of the block.

Modern models shift this further. In Llama 3 8B, d_model = 4096, the SwiGLU FFN has hidden size 14,336, and GQA shrinks the K and V projections:

Text
FFN       3 × 4096 × 14336                     ≈ 176M per layerattention 4096² (Q) + 4096² (O) + 2 × 4096 × 1024 (K, V) ≈ 42M per layerFFN share ≈ 81%

So in a 2026 LLM, the FFN is the dominant cost in parameters and in FLOPs at short context. It is also what Mixture-of-Experts models replicate: each "expert" is an FFN.

A real-life example

A company wants to host its own code-completion model on a single GPU and must cut the model's size. The engineer's first idea is to reduce attention heads. Using the numbers above, she sees that attention is under a fifth of the per-layer weights; halving it barely moves memory. Reducing d_ff from 14,336 to 11,008 saves far more.

They later move to a Mixture-of-Experts version: 8 FFN experts per layer, 2 active per token. Total parameters grow, but each token runs through only 2 FFNs, so per-token compute stays close to the dense model. Both decisions depend on knowing that the FFN is where the weights live.

Follow-up questions to expect

  • "Why expand to 4× and come back?" — The wider hidden layer gives many more non-linear feature detectors; projecting back keeps the residual stream width fixed.
  • "Does the FFN see other tokens?" — No. Any context it uses was already mixed into the vector by attention.
  • "What replaced ReLU in the FFN?" — GELU in BERT and GPT-2; SwiGLU in most LLMs since 2023.