Deep Learning Essentials

Course Content

Deep Learning Essentials

13 sections · 61 lessons

Which activation function do we use for each layer?


What you need to know

Hidden layers: pick for good gradients

ActivationFormula (in words)Where it is used
ReLUmax(0, x)Default for MLPs and CNNs; fast, gradient 1 for positive inputs
LeakyReLUx if positive, else 0.01 × xWhen many ReLU units die (always output 0)
GELUx weighted by the probability a standard normal is below xBERT, GPT-style transformers, ViT
SwiGLUa gated unit using the Swish functionMany modern LLMs' feed-forward layers
Tanhsquashes to −1 to 1Inside LSTM and GRU cells
Sigmoidsquashes to 0 to 1Gates in LSTM and GRU ("how much to let through")

For hidden layers, the choice between ReLU, GELU and LeakyReLU rarely changes results a lot. Using sigmoid in deep hidden layers does, because of vanishing gradients.

Output layer: pick for the task

TaskOutput layerActivationPyTorch loss
Binary (spam or not)1 unitSigmoidnn.BCEWithLogitsLoss
Multi-class, one label (which of 20 categories)20 unitsSoftmaxnn.CrossEntropyLoss
Multi-label (several tags can be true)1 unit per tagSigmoid per unitnn.BCEWithLogitsLoss
Regression (price, ETA)1 unitNonenn.MSELoss or nn.L1Loss
Regression, must be positive1 unitSoftplus or expnn.MSELoss, or a log-scale target

Why the losses take logits. nn.CrossEntropyLoss applies log-softmax internally, and nn.BCEWithLogitsLoss applies sigmoid internally. Combining the activation with the log inside the loss is more numerically stable than computing a probability and then taking its log. So in PyTorch, the model's last layer is a plain nn.Linear, and you apply softmax or sigmoid only at prediction time:

Python
logits = model(x)                           # shape (batch, 20), no softmaxloss = nn.CrossEntropyLoss()(logits, y)     # y holds class indicesprobs = logits.softmax(dim=1)               # only when you need probabilities

Softmax makes classes compete: raising one probability lowers the others, since they sum to 1. Sigmoid treats each output separately, which is exactly what multi-label needs.

A real-life example

A fashion marketplace builds one model to tag each product photo. It has three output heads on a shared CNN backbone:

  • Category (kurta, saree, jeans... 20 options, exactly one true): 20 units, softmax, cross-entropy.
  • Attributes (cotton, full-sleeve, printed, festive... several can be true): 12 units, a sigmoid each, binary cross-entropy.
  • Suggested price (a number): 1 unit, no activation, trained on the log of the price with MSE so a ₹300 kurta and a ₹30,000 saree have comparable errors.

A first version used softmax for the attributes. The model could never say a kurta was both "cotton" and "printed" with high confidence, because softmax forced the 12 attribute probabilities to sum to 1. Switching that head to sigmoid fixed it.

Follow-up questions to expect

  • "Why not use ReLU on a regression output?" — ReLU outputs 0 for all negative pre-activations, and the gradient there is 0, so if the output unit goes negative it can get stuck predicting 0. Use no activation, or softplus if the value must be positive.
  • "What is GELU and why do transformers use it?" — A smooth version of ReLU that lets small negative values through slightly. It trains a little better in transformers in practice and has no hard kink.
  • "Can hidden layers use different activations?" — Yes, but there is rarely a reason to. Consistency makes the network easier to initialise and debug.