Course Content
Deep Learning with TensorFlow and PyTorch
4 sections · 15 lessons
Activation Functions (ReLU, Sigmoid, Tanh)
Run an experiment. Take the two-moons dataset — two interleaved crescent shapes, a classic non-linear classification problem — and train a network with four hidden layers of 64 units each. That is a substantial model, about 12,700 parameters for a problem with two input features. Then delete every activation function, so each layer is a bare matrix multiply.
1import torch, torch.nn as nn2from sklearn.datasets import make_moons34X, y = make_moons(n_samples=2000, noise=0.15, random_state=0)56no_activation = nn.Sequential(7 nn.Linear(2, 64), nn.Linear(64, 64), nn.Linear(64, 64),8 nn.Linear(64, 64), nn.Linear(64, 1),9)10with_relu = nn.Sequential(11 nn.Linear(2, 64), nn.ReLU(), nn.Linear(64, 64), nn.ReLU(),12 nn.Linear(64, 64), nn.ReLU(), nn.Linear(64, 64), nn.ReLU(),13 nn.Linear(64, 1),14)15# After 500 epochs of Adam at lr=0.01:16# no_activation -> ~88% accuracy, and it will never get better17# with_relu -> ~99.5% accuracyThe deep network without activations cannot beat a single line drawn through the data, because that is precisely what it is. Five stacked matrix multiplications W5W4W3W2W1x collapse algebraically into one matrix Wx. The 12,700 parameters are real, they are all being trained, and they are jointly expressing a function with the power of about three.
The activation function is the one component that stops this collapse. Everything else in a network — depth, width, clever initialisation, fancy optimisers — is irrelevant without it. So it is worth knowing precisely what each candidate does, and, more usefully, how each one fails.
What the job actually requires
An activation function is applied element by element to the pre-activation values coming out of a layer. To be a good one it needs several properties at once, and the history of the field is largely a story of discovering that the early choices satisfied some of them and quietly violated others.
| Requirement | Why it matters | What breaks without it |
|---|---|---|
| Non-linear | Prevents the layer collapse shown above | Deep network has the power of a linear model |
| Differentiable (almost everywhere) | Backpropagation needs a derivative to pass through | No gradient, no learning |
| Derivative not tiny over the usual input range | The error signal is multiplied by this derivative at every layer | Vanishing gradients; early layers stop learning |
| Cheap to compute | Applied to every unit, every example, every step | Training slows measurably |
| Roughly zero-centred output | Keeps the next layer's inputs balanced around zero | Gradients for a whole layer share a sign; optimisation zig-zags |
| Unbounded above (usually) | Avoids saturation when activations are large | Units get stuck at their ceiling and stop responding |
Sigmoid: historically first, now mostly wrong
It squashes any real number into (0,1), which made it attractive: the output looks like a probability, and it is a smooth version of the perceptron's hard threshold. For twenty years it was the default.
Now look at its derivative. The maximum of σ(1−σ) occurs at σ=0.5, giving 0.5×0.5=0.25. So the best case multiplier that a gradient encounters at each sigmoid layer is one quarter, and the typical case is worse.
| Input z | σ(z) | σ′(z) | Consequence for the gradient |
|---|---|---|---|
| 0 | 0.500 | 0.250 | Best possible: quartered |
| 2 | 0.881 | 0.105 | Cut to a tenth |
| 4 | 0.982 | 0.018 | Cut to a fiftieth |
| 8 | 0.9997 | 0.0003 | Effectively erased |
Stack ten sigmoid layers and, even in the ideal case, the gradient reaching layer one is 0.2510≈10−6 of the gradient at the output. In practice inputs are rarely sitting at exactly zero, so it is far worse. The early layers receive a signal that rounds to nothing and never train.
Sigmoid has a second, subtler flaw: its output is always positive. Every unit in the following layer therefore receives all-positive inputs, and since the weight gradient is incoming activation × error signal, every weight feeding a given unit gets a gradient with the same sign. The whole row of weights must move up together or down together, so the optimiser cannot move diagonally and instead zig-zags towards the minimum.
Sigmoid still has exactly one correct home: the final layer of a binary classifier, where you genuinely want a probability. Using it in hidden layers is a bug with a plausible-looking cause.
Tanh: sigmoid, recentred
Output range (−1,1), so it is zero-centred — the zig-zag problem goes away. Its maximum derivative is 1 rather than 0.25, so gradients survive four times better per layer. It is a strict improvement on sigmoid for hidden layers.
It still saturates, though. At z=3, tanh′(3)≈0.0099; the gradient is cut by a factor of a hundred. So tanh delays the vanishing-gradient problem rather than solving it. It remains genuinely useful in one place: the internal gates of recurrent architectures such as LSTMs, where the bounded output is doing real work in keeping a recurrent state from exploding over long sequences.
ReLU: the one that unlocked depth
It looks almost too crude to work. It has a corner at zero, it is unbounded above, and it throws away all negative information. It is also the single most important activation function in the field, for three reasons.
The derivative is exactly 1 for positive inputs. No shrinking factor. A gradient can pass back through fifty ReLU layers and arrive undiminished, provided the units are active. That fact alone made networks deeper than about eight layers trainable for the first time.
It is nearly free to compute. A comparison and a select, versus the exponentials that sigmoid and tanh require. On large models this is a measurable fraction of training time.
It produces sparsity. Roughly half the units output exactly zero for any given input, which makes the representation sparse and, empirically, more robust.
The failure mode: dead ReLUs
Because ReLU′(z)=0 for all negative z, a unit whose pre-activation is negative for every training example receives zero gradient forever. Its weights never change. It is dead.
Here is how a unit dies, with numbers. Suppose a unit has weights that produce z=0.3 on average, and one large gradient step (a learning rate of 0.5 applied to a gradient of 2.0, say) shifts its bias by −1.0. Now z=−0.7 for typical inputs. The ReLU outputs 0. The derivative is 0. The gradient flowing to that unit's weights is 0× anything =0. Nothing will ever push the bias back up. The unit is permanently a constant zero, and its parameters are wasted.
You can measure this directly:
1activations = {}23def record(name):4 def hook(module, inp, out):5 activations[name] = out.detach()6 return hook78for name, layer in model.named_modules():9 if isinstance(layer, nn.ReLU):10 layer.register_forward_hook(record(name))1112model(next(iter(train_loader))[0])13for name, act in activations.items():14 # fraction of units that are zero for EVERY example in the batch15 dead = (act == 0).all(dim=0).float().mean().item()16 print(f"{name}: {dead:.1%} of units never fire")Under 10% is normal and healthy. Over 40% means you are wasting most of the layer, and the usual culprit is a learning rate that is too high, or an initialisation that started the biases too negative.
The ReLU family: patching the dead-unit hole
| Function | Definition for z<0 | Gradient when z<0 | Cost | When to use |
|---|---|---|---|---|
| ReLU | 0 | 0 | Cheapest | Default for CNNs and MLPs |
| Leaky ReLU | αz, α=0.01 | 0.01 | Cheap | When you have measured dead units |
| PReLU | αz, α learned | learned | Cheap, adds parameters | Large datasets where a per-channel slope is worth learning |
| ELU | α(ez−1) | →0 smoothly | Exponential | When smoothness near zero helps optimisation |
| GELU | z⋅Φ(z) | smooth, small | Exponential | Transformers — the de facto standard there |
| SiLU / Swish | z⋅σ(z) | smooth, can be negative | Exponential | Modern vision backbones, often a small win over ReLU |
A useful piece of realism: switching from ReLU to Leaky ReLU on a healthy network typically changes final accuracy by a fraction of a percent. These variants are worth reaching for when you have diagnosed a problem, not as a default hyperparameter to sweep. The exception is GELU in transformer architectures, where it is simply the convention and departing from it costs you comparability with published results.
1import torch.nn as nn23nn.ReLU() # max(0, z)4nn.LeakyReLU(0.01) # small negative slope5nn.PReLU(num_parameters=64) # one learned slope per channel6nn.ELU(alpha=1.0)7nn.GELU()8nn.SiLU() # also called SwishSoftmax: not really an activation function
Softmax belongs in a different category. It does not act element by element; it looks at a whole vector of scores at once and turns it into a probability distribution.
Take three class scores z=[2.0, 1.0, 0.1]:
| Class | Score zi | ezi | Probability |
|---|---|---|---|
| A | 2.0 | 7.389 | 0.659 |
| B | 1.0 | 2.718 | 0.242 |
| C | 0.1 | 1.105 | 0.099 |
| Total | 11.212 | 1.000 |
Two behaviours to internalise. It is shift-invariant: adding the same constant to every score leaves the probabilities unchanged, because the constant factors out of numerator and denominator. And it is competitive: the outputs must sum to 1, so raising one class's probability necessarily lowers another's. That competition is correct when exactly one class is right and catastrophic when several can be right simultaneously — for multi-label problems you want an independent sigmoid on each output instead.
The overflow trap, and why you should not call softmax yourself
e1000 overflows to infinity in float32, and ∞/∞ is nan. Shift-invariance gives the standard fix — subtract the maximum score before exponentiating, which changes nothing mathematically and keeps every exponent at or below e0=1:
1import numpy as np23def softmax_naive(z):4 e = np.exp(z)5 return e / e.sum()67def softmax_stable(z):8 e = np.exp(z - z.max()) # largest exponent is now exactly 09 return e / e.sum()1011big = np.array([1000.0, 999.0, 998.0])12print(softmax_naive(big)) # [nan nan nan]13print(softmax_stable(big)) # [0.665 0.245 0.090]In real code you should not write either version. Both frameworks provide loss functions that take raw scores (logits) and do the stable softmax internally, in a fused operation with better numerics than doing the two steps separately:
1# PyTorch -- final layer outputs raw logits, no activation2model = nn.Sequential(nn.Linear(128, 10)) # NOT nn.Softmax at the end3loss_fn = nn.CrossEntropyLoss() # applies log-softmax itself4loss = loss_fn(model(x), targets) # targets are class indices56# Binary case: same principle7loss_fn = nn.BCEWithLogitsLoss() # NOT nn.Sigmoid + nn.BCELoss1# TensorFlow -- same idea, spelled with a flag2model.compile(3 loss=tf.keras.losses.SparseCategoricalCrossentropy(from_logits=True),4 optimizer="adam",5)Applying softmax in the model and using a loss that expects logits is one of the most common silent bugs in the field. The model still trains; it just trains badly, because the softmax has already flattened the score differences the loss was designed to exploit. There is no error message. You only notice because accuracy sits several points below where it should.
Choosing, in practice
| Where | Use | Reason |
|---|---|---|
| Hidden layers, CNN or MLP | ReLU | Fast, gradient-preserving, well understood. Start here every time. |
| Hidden layers, transformer | GELU | Convention, and matched to published architectures |
| Hidden layers, measured > 40% dead units | Leaky ReLU (0.01) | Keeps a trickle of gradient flowing to inactive units |
| Recurrent gates (LSTM, GRU) | tanh and sigmoid | Bounded outputs keep recurrent state stable |
| Output, binary classification | Sigmoid — via BCEWithLogitsLoss | Single probability |
| Output, single-label multi-class | Softmax — via CrossEntropyLoss | Competing probabilities summing to 1 |
| Output, multi-label | Independent sigmoids | Labels must not compete for a shared probability budget |
| Output, regression | None | Any real value must be reachable |
When something is going wrong, the activation function is worth suspecting in a specific set of circumstances rather than as a general first move. If the loss plateaus at a high value and the gradient norms in early layers are around 10−7, you have vanishing gradients — check whether sigmoid or tanh has crept into a hidden layer. If a large fraction of units output zero for every input in a batch, you have dead ReLUs — lower the learning rate first, and only then reach for Leaky ReLU. If accuracy is a few points below a published baseline for no visible reason, check whether your final layer has a softmax on it that your loss function is also applying.
And if a deep model performs no better than logistic regression, count your activation functions. There should be one between every consecutive pair of linear layers, and none after the last one.