Course Content
Deep Learning Essentials
13 sections · 61 lessons
What are forward and backward propagation?
What you need to know
The forward pass: turning input into a loss
Each layer takes the previous layer's output a_prev and computes z = W·a_prev + b, then a = activation(z). The last layer's output is the prediction. A loss function then compares the prediction with the true label and returns one number: how wrong the network was.
During training, the framework also stores every intermediate z and a. It needs them for the backward pass. This is why training uses much more GPU memory than inference.
The backward pass: sharing out the blame
The loss depends on the last layer, which depends on the layer before it, and so on. The chain rule says you can get the gradient for an early weight by multiplying the local derivatives along the path back from the loss. Backpropagation does this in reverse order, so each layer reuses the gradient already computed for the layer after it.
That reuse is what makes it fast. Computing all gradients costs roughly the same as a couple of forward passes, no matter how many weights there are. The naive alternative, nudging each weight separately and re-running the network, would need one forward pass per weight, which is impossible for a model with millions of weights.
Watch the numbers on one weight
Take one neuron with no activation: pred = w × x. Let x = 2, w = 0.5, true value y = 3, loss = (pred − y)².
- Forward —
pred = 0.5 × 2 = 1.0,loss = (1 − 3)² = 4.0. - Backward —
∂loss/∂pred = 2 × (pred − y) = −4, and∂pred/∂w = x = 2, so by the chain rule∂loss/∂w = −4 × 2 = −8. - Update — with learning rate 0.1:
w = 0.5 − 0.1 × (−8) = 1.3. - Check — new
pred = 2.6, newloss = 0.16. One step cut the loss by 96%.
PyTorch does the backward step for you:
1import torch23x, y = torch.tensor(2.0), torch.tensor(3.0)4w = torch.tensor(0.5, requires_grad=True)56pred = w * x # forward7loss = (pred - y) ** 2 # 4.08loss.backward() # backward: fills w.grad9print(w.grad) # tensor(-8.)requires_grad=True tells autograd to record the operations on w. loss.backward() walks that record in reverse, applying the chain rule, and stores the result in w.grad. It does not change w; the optimiser does that.
A real-life example
A food-delivery app predicts delivery time from distance, restaurant prep time, rain and time of day. For one order the forward pass predicts 32 minutes; the order actually took 40. The loss is (32 − 40)² = 64.
The backward pass then answers: which weights caused the 8-minute miss, and by how much? Suppose the order was in heavy rain. The weight on the "rain" feature gets a large negative gradient, because a bigger rain weight would have pushed the prediction up towards 40. The weight on "time of day" gets a small gradient, because that feature barely moved the prediction. After millions of orders, these small corrections add up to a model that has learned how much rain really slows riders.
Follow-up questions to expect
- "Why does training need more memory than inference?" — The forward pass must keep every layer's activations so the backward pass can use them. Inference throws them away as soon as the next layer is computed.
- "Is backpropagation the same as gradient descent?" — No. Backpropagation computes the gradients; gradient descent (or Adam, or any optimiser) uses them to update the weights.
- "What is autograd?" — The part of PyTorch that records operations on tensors during the forward pass and runs the chain rule backwards over that record when you call
backward().