Course Content
Deep Learning Essentials
13 sections · 61 lessons
Why does tanh help a neural network converge faster than the logistic sigmoid?
What you need to know
How the two functions relate
Tanh is a rescaled sigmoid: tanh(x) = 2 × σ(2x) − 1. Same S-shape, but stretched to the range −1 to 1 and centred on 0.
| Sigmoid | Tanh | |
|---|---|---|
| Output range | 0 to 1 | −1 to 1 |
| Output at x = 0 | 0.5 | 0 |
| Maximum derivative | 0.25 | 1 |
| Zero-centred | No | Yes |
| Saturates at extremes | Yes | Yes |
Why "zero-centred" matters: the sign problem
Take a neuron in layer 2 with weights w1, w2, w3. Its inputs are layer 1's outputs a1, a2, a3. The gradient for each weight is:
∂Loss/∂w_i = δ × a_iwhere δ is the same error signal for the whole neuron. If layer 1 uses sigmoid, every a_i is positive, so every ∂Loss/∂w_i has the same sign as δ. In one step, all three weights must go up together or down together.
Suppose the best move is "increase w1, decrease w2". The optimiser cannot do that in one step. It has to alternate: all up, then all down, then all up, drifting slowly in the right direction. That zig-zag wastes steps. With tanh, the a_i have mixed signs, so the gradients can too, and the weights can move in the direction they actually need.
This is the same reason we standardise input features to mean zero: it removes the same zig-zag at the first layer.
Why the steeper slope matters
At the centre, tanh's derivative is 1 and sigmoid's is 0.25. So each tanh layer passes back up to four times more gradient than a sigmoid layer. Over several layers, that is the difference between early layers learning slowly and barely learning at all.
Where this leaves tanh today
Tanh still saturates for large inputs, with a derivative near 0.0099 at x = 3. So for deep hidden layers, ReLU-family activations beat both. Tanh remains standard where values must be bounded and signed: the candidate values and output of LSTM and GRU cells, and some output layers.
A real-life example
A factory monitors motors with a small 3-layer network on vibration-sensor features to predict failure within a week. With sigmoid hidden layers, validation loss takes about 60 epochs to settle. The engineer swaps them for tanh, keeping everything else the same, and the loss settles in about 25 epochs. Plotting two weights of one neuron over training shows why: under sigmoid they move in a staircase, both up then both down; under tanh they move almost diagonally.
He then tries ReLU with He initialisation, which settles slightly faster still. The team keeps ReLU in the hidden layers and a sigmoid only at the output, for the failure probability.
Follow-up questions to expect
- "Isn't sigmoid still needed for probabilities?" — Yes, at the output layer. Zero-centring matters for hidden layers, whose outputs feed other layers.
- "Does batch normalisation remove the difference?" — Largely, because it re-centres activations before the next layer. That is one reason the choice matters less in modern networks.
- "Why does an LSTM use tanh for its cell output?" — To keep values in a bounded, signed range so they can both increase and decrease the state, while sigmoid gates decide how much of it to pass.