Course Content
Deep Learning Essentials
13 sections · 61 lessons
What is an activation function, and why is it needed?
What you need to know
What a neuron does without and with an activation
A neuron computes a weighted sum z = w1·x1 + w2·x2 + ... + b. That is a linear function of the inputs: double the inputs and z doubles (apart from the bias). An activation a = f(z) then transforms z. If f is non-linear, such as ReLU, sigmoid or tanh, the neuron's output is no longer a straight-line function of its input.
Why linear layers collapse
Two linear layers in a row:
layer 1: h = W1·x + b1layer 2: y = W2·h + b2 = W2·(W1·x + b1) + b2 = (W2·W1)·x + (W2·b1 + b2) = W·x + b # one linear layerThe product of two matrices is just another matrix. So a 50-layer network with no activations has exactly the power of logistic or linear regression. You can check this in PyTorch:
1import torch, torch.nn as nn23net = nn.Sequential(nn.Linear(4, 8), nn.Linear(8, 3)) # no activation4W = net[1].weight @ net[0].weight5b = net[1].weight @ net[0].bias + net[1].bias6x = torch.randn(5, 4)7print(torch.allclose(net(x), x @ W.T + b, atol=1e-6)) # TrueThe two layers' weights multiply into one matrix W, and the network's output equals a single linear layer's output for any input. Insert nn.ReLU() between them and the equality breaks.
What non-linearity buys you
The classic example is XOR: output 1 when exactly one of two inputs is 1. No single straight line separates the 1s from the 0s. A network with one hidden ReLU layer of two units solves it, because ReLU lets the hidden layer fold the input space so the classes become separable.
More generally, each ReLU unit is "off" on one side of a line and "on" on the other. Combining many such units builds boundaries out of many straight pieces, which can approximate any curve.
What makes a good activation
- Non-linear, or the collapse above happens.
- Differentiable almost everywhere, so backpropagation can compute gradients. ReLU has a kink at 0, and frameworks simply use 0 or 1 there.
- Cheap, since it runs on every neuron for every example.
- Gradient that does not vanish: its derivative should not be tiny across most inputs. This is why ReLU replaced sigmoid in hidden layers.
A real-life example
A lender predicts loan default from income, existing EMIs and credit history. Risk is not a straight line: it stays low while EMIs are below about 40% of income, then rises steeply. A linear model can only say "each extra 1% of EMI adds the same risk", so it overestimates risk for low-EMI borrowers and underestimates it for high-EMI ones.
A small network with ReLU hidden units learns the bend: some units stay at zero until the EMI ratio crosses a point, then switch on and add risk quickly. The engineer notices that removing the nn.ReLU() lines by mistake, during a refactor, makes the network's accuracy drop to exactly that of logistic regression, which is the collapse above showing up in production.
Follow-up questions to expect
- "Is ReLU really non-linear? It is linear on each side." — Yes. It is piecewise linear, and the kink at 0 is enough. Many ReLUs together form any piecewise-linear shape, which can approximate smooth curves.
- "Does the output layer need an activation?" — Only to match the task: sigmoid for a probability, softmax for class probabilities, none for regression.
- "What is the universal approximation theorem?" — It says a network with one hidden layer and a non-linear activation can approximate any continuous function on a bounded range, given enough units. It says nothing about how easy that is to train, which is why depth helps in practice.