Course Content
Deep Learning Essentials
13 sections · 61 lessons
Which activation function do we use for each layer?
What you need to know
Hidden layers: pick for good gradients
| Activation | Formula (in words) | Where it is used |
|---|---|---|
| ReLU | max(0, x) | Default for MLPs and CNNs; fast, gradient 1 for positive inputs |
| LeakyReLU | x if positive, else 0.01 × x | When many ReLU units die (always output 0) |
| GELU | x weighted by the probability a standard normal is below x | BERT, GPT-style transformers, ViT |
| SwiGLU | a gated unit using the Swish function | Many modern LLMs' feed-forward layers |
| Tanh | squashes to −1 to 1 | Inside LSTM and GRU cells |
| Sigmoid | squashes to 0 to 1 | Gates in LSTM and GRU ("how much to let through") |
For hidden layers, the choice between ReLU, GELU and LeakyReLU rarely changes results a lot. Using sigmoid in deep hidden layers does, because of vanishing gradients.
Output layer: pick for the task
| Task | Output layer | Activation | PyTorch loss |
|---|---|---|---|
| Binary (spam or not) | 1 unit | Sigmoid | nn.BCEWithLogitsLoss |
| Multi-class, one label (which of 20 categories) | 20 units | Softmax | nn.CrossEntropyLoss |
| Multi-label (several tags can be true) | 1 unit per tag | Sigmoid per unit | nn.BCEWithLogitsLoss |
| Regression (price, ETA) | 1 unit | None | nn.MSELoss or nn.L1Loss |
| Regression, must be positive | 1 unit | Softplus or exp | nn.MSELoss, or a log-scale target |
Why the losses take logits. nn.CrossEntropyLoss applies log-softmax internally, and nn.BCEWithLogitsLoss applies sigmoid internally. Combining the activation with the log inside the loss is more numerically stable than computing a probability and then taking its log. So in PyTorch, the model's last layer is a plain nn.Linear, and you apply softmax or sigmoid only at prediction time:
logits = model(x) # shape (batch, 20), no softmaxloss = nn.CrossEntropyLoss()(logits, y) # y holds class indicesprobs = logits.softmax(dim=1) # only when you need probabilitiesSoftmax makes classes compete: raising one probability lowers the others, since they sum to 1. Sigmoid treats each output separately, which is exactly what multi-label needs.
A real-life example
A fashion marketplace builds one model to tag each product photo. It has three output heads on a shared CNN backbone:
- Category (kurta, saree, jeans... 20 options, exactly one true): 20 units, softmax, cross-entropy.
- Attributes (cotton, full-sleeve, printed, festive... several can be true): 12 units, a sigmoid each, binary cross-entropy.
- Suggested price (a number): 1 unit, no activation, trained on the log of the price with MSE so a ₹300 kurta and a ₹30,000 saree have comparable errors.
A first version used softmax for the attributes. The model could never say a kurta was both "cotton" and "printed" with high confidence, because softmax forced the 12 attribute probabilities to sum to 1. Switching that head to sigmoid fixed it.
Follow-up questions to expect
- "Why not use ReLU on a regression output?" — ReLU outputs 0 for all negative pre-activations, and the gradient there is 0, so if the output unit goes negative it can get stuck predicting 0. Use no activation, or softplus if the value must be positive.
- "What is GELU and why do transformers use it?" — A smooth version of ReLU that lets small negative values through slightly. It trains a little better in transformers in practice and has no hard kink.
- "Can hidden layers use different activations?" — Yes, but there is rarely a reason to. Consistency makes the network easier to initialise and debug.