Course Content
Transformer Architecture Q&A
6 sections · 60 lessons
What is SwiGLU, and why do modern LLMs prefer it over GELU/ReLU?
What you need to know
SiLU(x) = x · sigmoid(x) (Swish with beta = 1)SwiGLU(x) = W_down · ( SiLU(W_gate x) ⊙ (W_up x) ) ⊙ = element-wise multiplyThree matrices instead of two. To keep the parameter count the same as a 4 × d_model ReLU or GELU FFN, the hidden width is reduced to about 8/3 × d_model:
standard FFN : 2 matrices × d × 4d = 8 d²SwiGLU : 3 matrices × d × (8/3)d = 8 d²Real models round the width to a hardware-friendly multiple, and some go wider: Llama 2 7B uses 11,008 for d_model = 4096 (about 8/3), while Llama 3 8B uses 14,336 (3.5×).
Worked by hand: three hidden units
gate pre-activation W_gate x = [ 2.0, -1.0, 0.0]content W_up x = [ 0.5, 3.0, 4.0]SiLU(gate) = [ 1.762, -0.269, 0.0]gate ⊙ content = [ 0.881, -0.807, 0.0]Unit 3 has a large content value (4.0) but its gate is 0, so it is switched off for this token. Unit 2's gate is negative, so it flips and shrinks its content. The input decides, per unit, how much content flows through — that is the "gated linear unit" idea.
1import torch2from torch import nn3import torch.nn.functional as F45class SwiGLU(nn.Module):6 def __init__(self, d_model=4096, d_ff=14336): # Llama 3 8B sizes7 super().__init__()8 self.gate = nn.Linear(d_model, d_ff, bias=False)9 self.up = nn.Linear(d_model, d_ff, bias=False)10 self.down = nn.Linear(d_ff, d_model, bias=False)1112 def forward(self, x):13 return self.down(F.silu(self.gate(x)) * self.up(x)) # gate times content1415m = SwiGLU(d_model=64, d_ff=172) # tiny version to run on a laptop16print(m(torch.randn(1, 3, 64)).shape, sum(p.numel() for p in m.parameters()))17# torch.Size([1, 3, 64]) 33024The tiny version has 33,024 weights, close to the 32,768 of a standard FFN with d_ff = 256 — the 8/3 rule at work.
Why it is preferred, and what it costs
- Better quality at the same size. Shazeer's "GLU Variants Improve Transformer" (2020) compared variants at matched parameters and compute; gated versions, SwiGLU and GeGLU among them, gave lower perplexity, and the result has held up in large models.
- No clean theory. The paper itself says it offers no explanation and credits the gains to "divine benevolence".
- Costs. Three matmuls instead of two, an extra hidden-size tensor to store for backprop, and odd hidden sizes. These are small next to the quality gain.
A real-life example
A team designs a 1-billion-parameter code-completion model from scratch. They start with a GPT-2 style FFN (GELU, d_ff = 4 × 2048 = 8192). Following current practice, they switch to SwiGLU with d_ff = 5,632 (8/3 × 2048 ≈ 5,461, rounded up to a multiple of 256 for efficient GPU kernels). The parameter count stays nearly the same.
In their small-scale comparison on held-out code, the SwiGLU run reaches the same validation loss in fewer steps. The rounding matters too: an unrounded width like 5,461 leaves GPU tensor cores partly idle, so they always pick multiples of 64 or 256.
Follow-up questions to expect
- "What is GeGLU?" — The same gated design with GELU as the gate activation instead of SiLU; it performs similarly, and Google's Gemma models use it.
- "Why no biases?" — Most modern LLMs drop biases in linear layers; they add little and slightly complicate training and quantisation.
- "Does SwiGLU change the attention part?" — No, only the FFN.