Deep Learning Essentials

Course Content

Deep Learning Essentials

13 sections · 61 lessons

Why don't we use sigmoid or tanh in the hidden layers of a neural network?


Sigmoid's derivative across its input range0.00660.1050.250.1050.006601234x = −5x = 0,the peakx = 5Even the peak passes back only a quarter of the gradient.
Saturated sigmoid units pass back almost nothing, and the best case still shrinks the signal fourfold per layer.

What you need to know

Saturation, with numbers

The sigmoid is σ(x) = 1 / (1 + e^(−x)), and its derivative is σ(x) × (1 − σ(x)):

Input xσ(x)Derivative
00.500.25 (the maximum)
20.880.105
50.9930.0066

Tanh does better at the centre (derivative 1 at 0) but also saturates: at x = 3 its derivative is about 0.0099.

What that does across layers

In backpropagation, the gradient reaching layer 1 includes one activation derivative per layer. With sigmoid, even in the best case each layer multiplies by 0.25. Across 10 layers that is 0.25^10, about one-millionth. If some neurons are saturated, it is far worse. The last layers learn while the first layers, which learn the basic features, stay close to random.

Two smaller problems

  • Not zero-centred (sigmoid only). Sigmoid outputs are always positive, so every weight into the next neuron gets a gradient with the same sign. The weights can only all rise or all fall together in one step, which forces a zig-zag path. Tanh does not have this problem.
  • Cost. Both need an exponential. ReLU is a comparison with zero. For a large CNN this adds up.

Where sigmoid and tanh are still the right choice

  • Output layers: sigmoid for a probability.
  • Gates: LSTM and GRU gates use sigmoid because a gate must be between 0 (block) and 1 (pass). The cell's candidate values use tanh to keep them between −1 and 1.
  • Shallow networks: with two or three layers, saturation matters much less.
  • Bounded outputs: tanh when a value must stay in −1 to 1, such as some control or generative-model outputs.

A real-life example

A credit-card issuer's old risk model is a 12-layer MLP with sigmoid in every hidden layer. Training loss sits at about 0.69 for many epochs, which is the binary cross-entropy of a model that always predicts 50%, meaning it has learned almost nothing. Gradient norms show the last layer at around 1e-2 and the first layer at around 1e-9.

The engineer swaps the hidden sigmoids for ReLU, switches initialisation to He, and keeps a single sigmoid at the output for the default probability. The loss starts falling in the first epoch. The model's size and data did not change; only the activation's derivative did.

Follow-up questions to expect

  • "Would batch normalisation fix sigmoid networks?" — It helps, by keeping inputs to the sigmoid near 0 where the slope is largest, but the 0.25 maximum still shrinks gradients per layer. ReLU is simpler.
  • "Why is the sigmoid's maximum derivative 0.25?" — The derivative is σ × (1 − σ), which is largest when σ = 0.5, giving 0.5 × 0.5 = 0.25.
  • "Does ReLU have its own problems?" — Yes: units can die, outputting zero for every input. LeakyReLU and GELU reduce that.