Deep Learning with TensorFlow and PyTorch

Activation Functions (ReLU, Sigmoid, Tanh)


Run an experiment. Take the two-moons dataset — two interleaved crescent shapes, a classic non-linear classification problem — and train a network with four hidden layers of 64 units each. That is a substantial model, about 12,700 parameters for a problem with two input features. Then delete every activation function, so each layer is a bare matrix multiply.

Python
import torch, torch.nn as nnfrom sklearn.datasets import make_moonsX, y = make_moons(n_samples=2000, noise=0.15, random_state=0)no_activation = nn.Sequential(    nn.Linear(2, 64), nn.Linear(64, 64), nn.Linear(64, 64),    nn.Linear(64, 64), nn.Linear(64, 1),)with_relu = nn.Sequential(    nn.Linear(2, 64), nn.ReLU(), nn.Linear(64, 64), nn.ReLU(),    nn.Linear(64, 64), nn.ReLU(), nn.Linear(64, 64), nn.ReLU(),    nn.Linear(64, 1),)# After 500 epochs of Adam at lr=0.01:#   no_activation  -> ~88% accuracy, and it will never get better#   with_relu      -> ~99.5% accuracy

The deep network without activations cannot beat a single line drawn through the data, because that is precisely what it is. Five stacked matrix multiplications W5W4W3W2W1xW_5W_4W_3W_2W_1x collapse algebraically into one matrix WxWx. The 12,700 parameters are real, they are all being trained, and they are jointly expressing a function with the power of about three.

The activation function is the one component that stops this collapse. Everything else in a network — depth, width, clever initialisation, fancy optimisers — is irrelevant without it. So it is worth knowing precisely what each candidate does, and, more usefully, how each one fails.

Why depth waited for a non-saturating unitSigmoid and tanh saturate• Slope near zero once the input is large• Five layers multiply five tiny slopes• Sigmoid is not centred on zero• Gradient arrives at layer one as noiseReLU passes gradient through• Slope is exactly 1 on the positive side• No shrinkage no matter how deep• Cheap — a comparison, not an exponential• Cost: a unit stuck negative is dead
Saturation is a gradient problem, not an output problem: the forward values look fine while learning stops.

What the job actually requires

An activation function is applied element by element to the pre-activation values coming out of a layer. To be a good one it needs several properties at once, and the history of the field is largely a story of discovering that the early choices satisfied some of them and quietly violated others.

RequirementWhy it mattersWhat breaks without it
Non-linearPrevents the layer collapse shown aboveDeep network has the power of a linear model
Differentiable (almost everywhere)Backpropagation needs a derivative to pass throughNo gradient, no learning
Derivative not tiny over the usual input rangeThe error signal is multiplied by this derivative at every layerVanishing gradients; early layers stop learning
Cheap to computeApplied to every unit, every example, every stepTraining slows measurably
Roughly zero-centred outputKeeps the next layer's inputs balanced around zeroGradients for a whole layer share a sign; optimisation zig-zags
Unbounded above (usually)Avoids saturation when activations are largeUnits get stuck at their ceiling and stop responding

Sigmoid: historically first, now mostly wrong

σ(z)=11+e−zσ′(z)=σ(z)(1−σ(z))\sigma(z) = \frac{1}{1 + e^{-z}} \qquad \sigma'(z) = \sigma(z)\big(1 - \sigma(z)\big)

It squashes any real number into (0,1)(0, 1), which made it attractive: the output looks like a probability, and it is a smooth version of the perceptron's hard threshold. For twenty years it was the default.

Now look at its derivative. The maximum of σ(1−σ)\sigma(1-\sigma) occurs at σ=0.5\sigma = 0.5, giving 0.5×0.5=0.250.5 \times 0.5 = 0.25. So the best case multiplier that a gradient encounters at each sigmoid layer is one quarter, and the typical case is worse.

Input zzσ(z)\sigma(z)σ′(z)\sigma'(z)Consequence for the gradient
00.5000.250Best possible: quartered
20.8810.105Cut to a tenth
40.9820.018Cut to a fiftieth
80.99970.0003Effectively erased

Stack ten sigmoid layers and, even in the ideal case, the gradient reaching layer one is 0.2510≈10−60.25^{10} \approx 10^{-6} of the gradient at the output. In practice inputs are rarely sitting at exactly zero, so it is far worse. The early layers receive a signal that rounds to nothing and never train.

Sigmoid has a second, subtler flaw: its output is always positive. Every unit in the following layer therefore receives all-positive inputs, and since the weight gradient is incoming activation × error signal, every weight feeding a given unit gets a gradient with the same sign. The whole row of weights must move up together or down together, so the optimiser cannot move diagonally and instead zig-zags towards the minimum.

Sigmoid still has exactly one correct home: the final layer of a binary classifier, where you genuinely want a probability. Using it in hidden layers is a bug with a plausible-looking cause.

Tanh: sigmoid, recentred

tanh⁡(z)=ez−e−zez+e−ztanh⁡′(z)=1−tanh⁡2(z)\tanh(z) = \frac{e^{z} - e^{-z}}{e^{z} + e^{-z}} \qquad \tanh'(z) = 1 - \tanh^2(z)

Output range (−1,1)(-1, 1), so it is zero-centred — the zig-zag problem goes away. Its maximum derivative is 1 rather than 0.25, so gradients survive four times better per layer. It is a strict improvement on sigmoid for hidden layers.

It still saturates, though. At z=3z = 3, tanh⁡′(3)≈0.0099\tanh'(3) \approx 0.0099; the gradient is cut by a factor of a hundred. So tanh delays the vanishing-gradient problem rather than solving it. It remains genuinely useful in one place: the internal gates of recurrent architectures such as LSTMs, where the bounded output is doing real work in keeping a recurrent state from exploding over long sequences.

ReLU: the one that unlocked depth

ReLU(z)=max⁡(0,z)ReLU′(z)={1z>00z<0\text{ReLU}(z) = \max(0, z) \qquad \text{ReLU}'(z) = \begin{cases}1 & z > 0 \\ 0 & z < 0\end{cases}

It looks almost too crude to work. It has a corner at zero, it is unbounded above, and it throws away all negative information. It is also the single most important activation function in the field, for three reasons.

The derivative is exactly 1 for positive inputs. No shrinking factor. A gradient can pass back through fifty ReLU layers and arrive undiminished, provided the units are active. That fact alone made networks deeper than about eight layers trainable for the first time.

It is nearly free to compute. A comparison and a select, versus the exponentials that sigmoid and tanh require. On large models this is a measurable fraction of training time.

It produces sparsity. Roughly half the units output exactly zero for any given input, which makes the representation sparse and, empirically, more robust.

The failure mode: dead ReLUs

Because ReLU′(z)=0\text{ReLU}'(z) = 0 for all negative zz, a unit whose pre-activation is negative for every training example receives zero gradient forever. Its weights never change. It is dead.

Here is how a unit dies, with numbers. Suppose a unit has weights that produce z=0.3z = 0.3 on average, and one large gradient step (a learning rate of 0.5 applied to a gradient of 2.0, say) shifts its bias by −1.0-1.0. Now z=−0.7z = -0.7 for typical inputs. The ReLU outputs 0. The derivative is 0. The gradient flowing to that unit's weights is 0×0 \times anything =0= 0. Nothing will ever push the bias back up. The unit is permanently a constant zero, and its parameters are wasted.

You can measure this directly:

Python
activations = {}def record(name):    def hook(module, inp, out):        activations[name] = out.detach()    return hookfor name, layer in model.named_modules():    if isinstance(layer, nn.ReLU):        layer.register_forward_hook(record(name))model(next(iter(train_loader))[0])for name, act in activations.items():    # fraction of units that are zero for EVERY example in the batch    dead = (act == 0).all(dim=0).float().mean().item()    print(f"{name}: {dead:.1%} of units never fire")

Under 10% is normal and healthy. Over 40% means you are wasting most of the layer, and the usual culprit is a learning rate that is too high, or an initialisation that started the biases too negative.

The ReLU family: patching the dead-unit hole

FunctionDefinition for z<0z < 0Gradient when z<0z < 0CostWhen to use
ReLU000CheapestDefault for CNNs and MLPs
Leaky ReLUαz\alpha z, α=0.01\alpha = 0.010.01CheapWhen you have measured dead units
PReLUαz\alpha z, α\alpha learnedlearnedCheap, adds parametersLarge datasets where a per-channel slope is worth learning
ELUα(ez−1)\alpha(e^{z} - 1)→0\to 0 smoothlyExponentialWhen smoothness near zero helps optimisation
GELUz⋅Φ(z)z \cdot \Phi(z)smooth, smallExponentialTransformers — the de facto standard there
SiLU / Swishz⋅σ(z)z \cdot \sigma(z)smooth, can be negativeExponentialModern vision backbones, often a small win over ReLU

A useful piece of realism: switching from ReLU to Leaky ReLU on a healthy network typically changes final accuracy by a fraction of a percent. These variants are worth reaching for when you have diagnosed a problem, not as a default hyperparameter to sweep. The exception is GELU in transformer architectures, where it is simply the convention and departing from it costs you comparability with published results.

Python
import torch.nn as nnnn.ReLU()                    # max(0, z)nn.LeakyReLU(0.01)           # small negative slopenn.PReLU(num_parameters=64)  # one learned slope per channelnn.ELU(alpha=1.0)nn.GELU()nn.SiLU()                    # also called Swish

Softmax: not really an activation function

Softmax belongs in a different category. It does not act element by element; it looks at a whole vector of scores at once and turns it into a probability distribution.

softmax(z)i=ezi∑jezj\text{softmax}(z)_i = \frac{e^{z_i}}{\sum_{j} e^{z_j}}

Take three class scores z=[2.0, 1.0, 0.1]z = [2.0,\ 1.0,\ 0.1]:

ClassScore ziz_iezie^{z_i}Probability
A2.07.3890.659
B1.02.7180.242
C0.11.1050.099
Total11.2121.000

Two behaviours to internalise. It is shift-invariant: adding the same constant to every score leaves the probabilities unchanged, because the constant factors out of numerator and denominator. And it is competitive: the outputs must sum to 1, so raising one class's probability necessarily lowers another's. That competition is correct when exactly one class is right and catastrophic when several can be right simultaneously — for multi-label problems you want an independent sigmoid on each output instead.

The overflow trap, and why you should not call softmax yourself

e1000e^{1000} overflows to infinity in float32, and ∞/∞\infty / \infty is nan. Shift-invariance gives the standard fix — subtract the maximum score before exponentiating, which changes nothing mathematically and keeps every exponent at or below e0=1e^0 = 1:

Python
import numpy as npdef softmax_naive(z):    e = np.exp(z)    return e / e.sum()def softmax_stable(z):    e = np.exp(z - z.max())     # largest exponent is now exactly 0    return e / e.sum()big = np.array([1000.0, 999.0, 998.0])print(softmax_naive(big))    # [nan nan nan]print(softmax_stable(big))   # [0.665 0.245 0.090]

In real code you should not write either version. Both frameworks provide loss functions that take raw scores (logits) and do the stable softmax internally, in a fused operation with better numerics than doing the two steps separately:

Python
# PyTorch -- final layer outputs raw logits, no activationmodel = nn.Sequential(nn.Linear(128, 10))       # NOT nn.Softmax at the endloss_fn = nn.CrossEntropyLoss()                 # applies log-softmax itselfloss = loss_fn(model(x), targets)               # targets are class indices# Binary case: same principleloss_fn = nn.BCEWithLogitsLoss()                # NOT nn.Sigmoid + nn.BCELoss
Python
# TensorFlow -- same idea, spelled with a flagmodel.compile(    loss=tf.keras.losses.SparseCategoricalCrossentropy(from_logits=True),    optimizer="adam",)

Applying softmax in the model and using a loss that expects logits is one of the most common silent bugs in the field. The model still trains; it just trains badly, because the softmax has already flattened the score differences the loss was designed to exploit. There is no error message. You only notice because accuracy sits several points below where it should.

Choosing, in practice

WhereUseReason
Hidden layers, CNN or MLPReLUFast, gradient-preserving, well understood. Start here every time.
Hidden layers, transformerGELUConvention, and matched to published architectures
Hidden layers, measured > 40% dead unitsLeaky ReLU (0.01)Keeps a trickle of gradient flowing to inactive units
Recurrent gates (LSTM, GRU)tanh and sigmoidBounded outputs keep recurrent state stable
Output, binary classificationSigmoid — via BCEWithLogitsLossSingle probability
Output, single-label multi-classSoftmax — via CrossEntropyLossCompeting probabilities summing to 1
Output, multi-labelIndependent sigmoidsLabels must not compete for a shared probability budget
Output, regressionNoneAny real value must be reachable

When something is going wrong, the activation function is worth suspecting in a specific set of circumstances rather than as a general first move. If the loss plateaus at a high value and the gradient norms in early layers are around 10−710^{-7}, you have vanishing gradients — check whether sigmoid or tanh has crept into a hidden layer. If a large fraction of units output zero for every input in a batch, you have dead ReLUs — lower the learning rate first, and only then reach for Leaky ReLU. If accuracy is a few points below a published baseline for no visible reason, check whether your final layer has a softmax on it that your loss function is also applying.

And if a deep model performs no better than logistic regression, count your activation functions. There should be one between every consecutive pair of linear layers, and none after the last one.