Course Content
Deep Learning Essentials
13 sections · 61 lessons
Why don't we use sigmoid or tanh in the hidden layers of a neural network?
What you need to know
Saturation, with numbers
The sigmoid is σ(x) = 1 / (1 + e^(−x)), and its derivative is σ(x) × (1 − σ(x)):
| Input x | σ(x) | Derivative |
|---|---|---|
| 0 | 0.50 | 0.25 (the maximum) |
| 2 | 0.88 | 0.105 |
| 5 | 0.993 | 0.0066 |
Tanh does better at the centre (derivative 1 at 0) but also saturates: at x = 3 its derivative is about 0.0099.
What that does across layers
In backpropagation, the gradient reaching layer 1 includes one activation derivative per layer. With sigmoid, even in the best case each layer multiplies by 0.25. Across 10 layers that is 0.25^10, about one-millionth. If some neurons are saturated, it is far worse. The last layers learn while the first layers, which learn the basic features, stay close to random.
Two smaller problems
- Not zero-centred (sigmoid only). Sigmoid outputs are always positive, so every weight into the next neuron gets a gradient with the same sign. The weights can only all rise or all fall together in one step, which forces a zig-zag path. Tanh does not have this problem.
- Cost. Both need an exponential. ReLU is a comparison with zero. For a large CNN this adds up.
Where sigmoid and tanh are still the right choice
- Output layers: sigmoid for a probability.
- Gates: LSTM and GRU gates use sigmoid because a gate must be between 0 (block) and 1 (pass). The cell's candidate values use tanh to keep them between −1 and 1.
- Shallow networks: with two or three layers, saturation matters much less.
- Bounded outputs: tanh when a value must stay in −1 to 1, such as some control or generative-model outputs.
A real-life example
A credit-card issuer's old risk model is a 12-layer MLP with sigmoid in every hidden layer. Training loss sits at about 0.69 for many epochs, which is the binary cross-entropy of a model that always predicts 50%, meaning it has learned almost nothing. Gradient norms show the last layer at around 1e-2 and the first layer at around 1e-9.
The engineer swaps the hidden sigmoids for ReLU, switches initialisation to He, and keeps a single sigmoid at the output for the default probability. The loss starts falling in the first epoch. The model's size and data did not change; only the activation's derivative did.
Follow-up questions to expect
- "Would batch normalisation fix sigmoid networks?" — It helps, by keeping inputs to the sigmoid near 0 where the slope is largest, but the 0.25 maximum still shrinks gradients per layer. ReLU is simpler.
- "Why is the sigmoid's maximum derivative 0.25?" — The derivative is
σ × (1 − σ), which is largest whenσ = 0.5, giving0.5 × 0.5 = 0.25. - "Does ReLU have its own problems?" — Yes: units can die, outputting zero for every input. LeakyReLU and GELU reduce that.