Transformer Architecture Q&A

Course Content

Transformer Architecture Q&A

6 sections · 60 lessons

Why has GELU largely replaced ReLU in Transformer FFNs?


What you need to know

Text
ReLU(x) = max(0, x)GELU(x) = x · Φ(x) = x · 0.5 · (1 + erf(x / sqrt(2)))tanh approximation: 0.5 · x · (1 + tanh( sqrt(2/π) · (x + 0.044715 x³) ))

Φ(x) is the standard normal cumulative distribution: close to 0 for very negative x, 0.5 at x = 0, close to 1 for large x.

Worked values

xReLU(x)GELU(x)
-3.00-0.004
-1.00-0.159
-0.750-0.170 (its minimum)
-0.50-0.154
0.000
0.50.50.346
1.01.00.841
2.02.01.955

For large positive inputs GELU is almost x; for large negative inputs it is almost 0. Between about -2 and 2 it curves smoothly.

The three reasons

  • Gradient below zero. ReLU's gradient is exactly 0 for any negative input. A unit whose inputs are always negative stops learning — a "dead neuron". GELU's gradient at x = -0.5 is about 0.13, so the unit can recover.
  • Smoothness. ReLU has a sharp corner at 0; GELU has none. Smooth functions tend to make optimisation behave better.
  • Non-monotonic dip. GELU goes slightly negative (lowest value -0.17 near x = -0.75) before returning to 0. This adds a little expressive shape ReLU lacks.

Honest caveats

The improvement is empirical and small. GELU costs more than ReLU (an erf or tanh), which is why the tanh approximation is common; today fused kernels hide most of that cost. And GELU itself has been largely replaced in new LLMs by SwiGLU, a gated unit that did better in the same kind of comparison.

A real-life example

A team fine-tunes a BERT-based document classifier for tax forms. An engineer, trying to speed up CPU inference, replaces GELU with ReLU in the FFNs of the pretrained model and fine-tunes again. Accuracy drops noticeably.

The reason: the pretrained weights were learned with GELU. With ReLU, every negative pre-activation that GELU passed as a small negative value is now zero, so each layer's output shifts, and fine-tuning on a small dataset cannot fully re-adapt 12 layers. They revert to GELU and instead get their speed from quantisation and a fused GELU kernel. The rule: never swap the activation of a pretrained model; choose it when training from scratch.

Follow-up questions to expect

  • "What is Swish or SiLU?" — x · sigmoid(x), a close relative of GELU with a very similar shape; it is the activation inside SwiGLU.
  • "Why not keep ReLU for speed?" — On GPUs the activation is a tiny share of runtime compared with the matrix multiplications, so the quality gain wins.
  • "Is GELU related to dropout?" — Its paper motivates it as the expected value of randomly zeroing an input, with a probability that depends on the input's size.