Course Content
Transformer Architecture Q&A
6 sections · 60 lessons
Why has GELU largely replaced ReLU in Transformer FFNs?
What you need to know
ReLU(x) = max(0, x)GELU(x) = x · Φ(x) = x · 0.5 · (1 + erf(x / sqrt(2)))tanh approximation: 0.5 · x · (1 + tanh( sqrt(2/π) · (x + 0.044715 x³) ))Φ(x) is the standard normal cumulative distribution: close to 0 for very negative x, 0.5 at x = 0, close to 1 for large x.
Worked values
| x | ReLU(x) | GELU(x) |
|---|---|---|
| -3.0 | 0 | -0.004 |
| -1.0 | 0 | -0.159 |
| -0.75 | 0 | -0.170 (its minimum) |
| -0.5 | 0 | -0.154 |
| 0.0 | 0 | 0 |
| 0.5 | 0.5 | 0.346 |
| 1.0 | 1.0 | 0.841 |
| 2.0 | 2.0 | 1.955 |
For large positive inputs GELU is almost x; for large negative inputs it is almost 0. Between about -2 and 2 it curves smoothly.
The three reasons
- Gradient below zero. ReLU's gradient is exactly 0 for any negative input. A unit whose inputs are always negative stops learning — a "dead neuron". GELU's gradient at
x = -0.5is about 0.13, so the unit can recover. - Smoothness. ReLU has a sharp corner at 0; GELU has none. Smooth functions tend to make optimisation behave better.
- Non-monotonic dip. GELU goes slightly negative (lowest value -0.17 near
x = -0.75) before returning to 0. This adds a little expressive shape ReLU lacks.
Honest caveats
The improvement is empirical and small. GELU costs more than ReLU (an erf or tanh), which is why the tanh approximation is common; today fused kernels hide most of that cost. And GELU itself has been largely replaced in new LLMs by SwiGLU, a gated unit that did better in the same kind of comparison.
A real-life example
A team fine-tunes a BERT-based document classifier for tax forms. An engineer, trying to speed up CPU inference, replaces GELU with ReLU in the FFNs of the pretrained model and fine-tunes again. Accuracy drops noticeably.
The reason: the pretrained weights were learned with GELU. With ReLU, every negative pre-activation that GELU passed as a small negative value is now zero, so each layer's output shifts, and fine-tuning on a small dataset cannot fully re-adapt 12 layers. They revert to GELU and instead get their speed from quantisation and a fused GELU kernel. The rule: never swap the activation of a pretrained model; choose it when training from scratch.
Follow-up questions to expect
- "What is Swish or SiLU?" —
x · sigmoid(x), a close relative of GELU with a very similar shape; it is the activation inside SwiGLU. - "Why not keep ReLU for speed?" — On GPUs the activation is a tiny share of runtime compared with the matrix multiplications, so the quality gain wins.
- "Is GELU related to dropout?" — Its paper motivates it as the expected value of randomly zeroing an input, with a probability that depends on the input's size.