Course Content
Deep Learning Essentials
13 sections · 61 lessons
What is the difference between ReLU and LeakyReLU functions?
What you need to know
The two functions, side by side
ReLU(x) = x if x > 0, else 0LeakyReLU(x) = x if x > 0, else 0.01 × xReLU'(x) = 1 if x > 0, else 0LeakyReLU'(x) = 1 if x > 0, else 0.01For an input of −3, ReLU outputs 0 with gradient 0. LeakyReLU outputs −0.03 with gradient 0.01. For +3, both output 3 with gradient 1.
How a ReLU unit dies
A common path: a large learning rate produces one big update that pushes a unit's bias to, say, −5. If the weighted inputs never exceed 5, the unit is negative for every example. Zero output means zero gradient, so the bias never moves back. The unit is now dead weight. In bad cases, a large share of a layer dies and the network's effective size shrinks without any error message.
Measuring it
1with torch.no_grad():2 z = model.fc1(x_val) # pre-activations, shape (N, units)3 dead = (z <= 0).all(dim=0).float().mean() # units never positive on any row4print(f"dead units in fc1: {dead:.1%}")(z <= 0).all(dim=0) is True for units that never fire on the validation batch. A few percent is normal; a third or more is a problem.
The ReLU family
- LeakyReLU: fixed small slope, usually 0.01.
- PReLU: the negative slope is a learned parameter.
- ELU: smooth exponential curve for negatives, outputs closer to zero mean.
- GELU: smooth, slightly negative for small negative inputs; the transformer default.
In practice these usually change accuracy only a little. Dying units are more often fixed by a lower learning rate, better initialisation or batch normalisation than by switching the activation.
A real-life example
A food-delivery app trains a demand model with a 3-layer ReLU MLP. To speed things up, an engineer raises the learning rate from 0.001 to 0.05. Training loss drops fast for 200 steps, then flattens at a worse value than before. Running the snippet above shows 58% of the first hidden layer's units are dead.
Two changes fix it: the learning rate goes back down with a warmup, and the hidden activation becomes nn.LeakyReLU(0.01). After retraining, dead units fall to about 3% and the loss beats the original run. The team adds the dead-unit check to their training dashboard.
Follow-up questions to expect
- "What is the derivative of ReLU at exactly 0?" — Mathematically undefined; frameworks just use 0. In practice an input of exactly 0.0 is rare and it does not matter.
- "Why is 0.01 the usual leak?" — It is small enough to behave almost like ReLU but large enough to keep gradients alive. PReLU lets the network learn it instead.
- "Is ReLU's zero output ever useful?" — Yes. It makes activations sparse, since many units are exactly 0 for a given input, which is cheap and often helps features specialise.