Transformer Architecture Q&A

Course Content

Transformer Architecture Q&A

6 sections · 60 lessons

What is SwiGLU, and why do modern LLMs prefer it over GELU/ReLU?


SwiGLU at Llama 3 8B sizesx, 4,096 wideGate: SiLU(xW_gate), 14,336 wideContent: xW_up,14,336 wideMultiply gateand contentelement-wiseW_downback to 4,096Three matrices at 8/3 width cost the same as two at 4x.
A hidden unit with content 4.0 but gate 0 contributes nothing — the input decides, unit by unit, what flows through.

What you need to know

Text
SiLU(x)   = x · sigmoid(x)             (Swish with beta = 1)SwiGLU(x) = W_down · ( SiLU(W_gate x) ⊙ (W_up x) )      ⊙ = element-wise multiply

Three matrices instead of two. To keep the parameter count the same as a 4 × d_model ReLU or GELU FFN, the hidden width is reduced to about 8/3 × d_model:

Text
standard FFN : 2 matrices × d × 4d      = 8 d²SwiGLU       : 3 matrices × d × (8/3)d  = 8 d²

Real models round the width to a hardware-friendly multiple, and some go wider: Llama 2 7B uses 11,008 for d_model = 4096 (about 8/3), while Llama 3 8B uses 14,336 (3.5×).

Worked by hand: three hidden units

Text
gate pre-activation  W_gate x = [ 2.0, -1.0, 0.0]content              W_up x   = [ 0.5,  3.0, 4.0]SiLU(gate)                    = [ 1.762, -0.269, 0.0]gate ⊙ content                = [ 0.881, -0.807, 0.0]

Unit 3 has a large content value (4.0) but its gate is 0, so it is switched off for this token. Unit 2's gate is negative, so it flips and shrinks its content. The input decides, per unit, how much content flows through — that is the "gated linear unit" idea.

Python
import torchfrom torch import nnimport torch.nn.functional as Fclass SwiGLU(nn.Module):    def __init__(self, d_model=4096, d_ff=14336):   # Llama 3 8B sizes        super().__init__()        self.gate = nn.Linear(d_model, d_ff, bias=False)        self.up = nn.Linear(d_model, d_ff, bias=False)        self.down = nn.Linear(d_ff, d_model, bias=False)    def forward(self, x):        return self.down(F.silu(self.gate(x)) * self.up(x))   # gate times contentm = SwiGLU(d_model=64, d_ff=172)                   # tiny version to run on a laptopprint(m(torch.randn(1, 3, 64)).shape, sum(p.numel() for p in m.parameters()))# torch.Size([1, 3, 64]) 33024

The tiny version has 33,024 weights, close to the 32,768 of a standard FFN with d_ff = 256 — the 8/3 rule at work.

Why it is preferred, and what it costs

  • Better quality at the same size. Shazeer's "GLU Variants Improve Transformer" (2020) compared variants at matched parameters and compute; gated versions, SwiGLU and GeGLU among them, gave lower perplexity, and the result has held up in large models.
  • No clean theory. The paper itself says it offers no explanation and credits the gains to "divine benevolence".
  • Costs. Three matmuls instead of two, an extra hidden-size tensor to store for backprop, and odd hidden sizes. These are small next to the quality gain.

A real-life example

A team designs a 1-billion-parameter code-completion model from scratch. They start with a GPT-2 style FFN (GELU, d_ff = 4 × 2048 = 8192). Following current practice, they switch to SwiGLU with d_ff = 5,632 (8/3 × 2048 ≈ 5,461, rounded up to a multiple of 256 for efficient GPU kernels). The parameter count stays nearly the same.

In their small-scale comparison on held-out code, the SwiGLU run reaches the same validation loss in fewer steps. The rounding matters too: an unrounded width like 5,461 leaves GPU tensor cores partly idle, so they always pick multiples of 64 or 256.

Follow-up questions to expect

  • "What is GeGLU?" — The same gated design with GELU as the gate activation instead of SiLU; it performs similarly, and Google's Gemma models use it.
  • "Why no biases?" — Most modern LLMs drop biases in linear layers; they add little and slightly complicate training and quantisation.
  • "Does SwiGLU change the attention part?" — No, only the FFN.