Course Content
Mathematics for Machine Learning
5 sections · 13 lessons
Partial Derivatives and Chain Rule
You can differentiate x2 in your sleep. Now here is the thing you are actually asked to differentiate when you train a network. A weight W1 feeds into a linear combination, which feeds into a sigmoid, which feeds into another linear combination, which feeds into a loss. Written out in one line, with x and the target y as fixed data:
You want ∂L/∂W1. Try attacking it directly: expand the square, and you are staring at σ(W1x+b1)2 with W1 buried inside a sigmoid inside a square. There is no rule in the standard table for "derivative of a squared sigmoid of a linear function of the thing I care about". And this is a network with one hidden unit. A real one has fifty layers and the expression would not fit on a page.
The way out is not a bigger table of rules. It is to stop treating the expression as one object and start treating it as a pipeline — a sequence of simple steps, each of which you can differentiate — and then find the rule for stitching those small derivatives together. That rule is the chain rule, and it is the single mathematical idea that makes training deep models possible.
First, differentiating with several variables at once
When a function takes more than one input, "the derivative" splits into one derivative per input. A partial derivative ∂f/∂x asks how the output responds to a change in x alone, with every other input frozen. Mechanically, you differentiate as usual and treat all the other variables as if they were constants.
Take a function of three variables:
Differentiate with respect to x, freezing y and z. The term x2y gives 2xy. The term yz3 has no x in it, so it is a constant and gives 0. The term sin(xz) needs the chain rule — derivative of sine is cosine, times the derivative of the inside, which is z:
Now with respect to y. Here x2y is a constant times y, giving x2; yz3 gives z3; and sin(xz) has no y, so it vanishes:
And with respect to z: x2y vanishes, yz3 gives 3yz2, and sin(xz) gives xcos(xz):
Collect them into a single vector and you have the gradient, ∇f=[∂f/∂x,∂f/∂y,∂f/∂z]T. It points in the direction of steepest increase, and its length is how steep that increase is. When f is a loss and the variables are parameters, −∇f is the direction that reduces the loss fastest.
A partial derivative is one number: how much this output moves per unit of this input, everything else held still.
Second derivatives, and the matrix that holds them
Differentiate twice and you learn about curvature rather than slope. With several variables there are several ways to differentiate twice, and they all matter.
For f(x,y)=x2y3:
- ∂f/∂x=2xy3, so ∂2f/∂x2=2y3 — how the x-slope changes as you move in x.
- ∂f/∂y=3x2y2, so ∂2f/∂y2=6x2y.
- The mixed one: differentiate ∂f/∂x=2xy3 with respect to y to get 6xy2. Or differentiate ∂f/∂y=3x2y2 with respect to x to get 6xy2. Same answer.
That agreement is not a coincidence. For any function with continuous second derivatives, the order of differentiation does not matter: ∂2f/∂x∂y=∂2f/∂y∂x. This is Clairaut's theorem, and its practical consequence is that the matrix collecting all second derivatives is symmetric.
That matrix is the Hessian:
If the gradient tells you which way the surface tilts, the Hessian tells you how it bends — whether you are in a bowl, on a dome, or on a mountain pass. Because it is symmetric, its eigenvalues are real, and their signs classify the point completely.
Classifying a critical point, worked through
Take f(x,y)=x3+y3−3xy. Critical points are where the gradient vanishes:
Substituting the first into the second: x=(x2)2=x4, so x4−x=0, so x(x3−1)=0. That gives x=0 (hence y=0) and x=1 (hence y=1). Two critical points, both with zero gradient. The gradient alone cannot tell them apart.
The second derivatives are ∂2f/∂x2=6x, ∂2f/∂y2=6y, and ∂2f/∂x∂y=−3.
At (0,0): H=[0−3−30]. Its characteristic equation is λ2−9=0, so λ=+3 and λ=−3. Mixed signs: the surface curves up along one direction and down along another. This is a saddle.
At (1,1): H=[6−3−36]. Trace 12, determinant 36−9=27, so λ2−12λ+27=0 and λ=9,3. Both positive: the surface curves upward in every direction. This is a genuine local minimum, with f(1,1)=1+1−3=−1.
| Eigenvalues of H | Point type | What an optimiser experiences |
|---|---|---|
| All positive | Local minimum | Settles and stays |
| All negative | Local maximum | Never reached by descent |
| Mixed signs | Saddle | Slows to a crawl, then eventually escapes down a negative direction |
| Some zero | Degenerate / flat | Plateau; gradients are tiny and progress stalls |
| All positive, widely spread | Ill-conditioned minimum | Zig-zags across the steep directions, crawls along the shallow ones |
That last row is the everyday one. It is not that the optimiser cannot find the minimum; it is that the ratio of the largest to smallest eigenvalue — the condition number — sets how many steps it takes. A ratio of 1000 means roughly a thousand-fold slowdown compared to a perfectly round bowl. Nobody computes the Hessian of a large network (it has one entry per pair of parameters, which for a million parameters is 1012 numbers), but this is the geometry that momentum and adaptive methods exist to compensate for.
The chain rule: differentiating a pipeline
If y depends on u, and u depends on x, then
The plain-English version is an exchange-rate argument. Suppose a 1-unit change in x produces a 3-unit change in u, and a 1-unit change in u produces a 2-unit change in y. Then a 1-unit change in x produces a 3×2=6-unit change in y. Rates multiply along a chain.
It extends to any depth. For y=sin(u), u=v2, v=3x+1:
At x=0: v=1, u=1, so the derivative is 6×1×cos(1)=6×0.5403=3.242. Notice that to evaluate the derivative you needed the intermediate values v and u from the forward computation. Hold on to that observation — it is the reason backpropagation has the memory cost it does.
When the intermediate values are several
If f depends on several variables and each of those depends on a common parameter t, every path contributes and the contributions add:
Take f(x,y)=x2+y2 with x=t2 and y=t3. Through the chain rule:
Check it by substituting first: f=t4+t6, whose derivative is 4t3+6t5. Identical. At t=2 that is 32+192=224.
Multiply along a path; add across paths. That one sentence is the entire chain rule, in any number of dimensions.
When the intermediates are vectors rather than single numbers, the per-variable partials get organised into a matrix called the Jacobian, and the chain rule becomes a matrix product instead of a scalar product. The logic is unchanged.
Backpropagation is the chain rule, applied once, carefully
Here is the network from the opening, with concrete numbers. One input, one hidden unit with a sigmoid, one linear output, squared-error loss.
x = 1.0 W1 = 0.5 b1 = 0.1y = 1.0 W2 = 0.8 b2 = -0.2z1 = W1*x + b1 a1 = sigmoid(z1)z2 = W2*a1 + b2 yhat = z2L = (yhat - y)^2Forward pass
| Quantity | Computation | Value |
|---|---|---|
| z1 | 0.5×1.0+0.1 | 0.6000 |
| a1 | σ(0.6)=1/(1+e−0.6) | 0.6457 |
| z2 | 0.8×0.6457−0.2 | 0.3165 |
| y^ | z2 | 0.3165 |
| L | (0.3165−1.0)2 | 0.4672 |
Backward pass
Now walk backwards, computing the derivative of L with respect to each intermediate, and multiplying by the local derivative at each step.
Output. ∂L/∂y^=2(y^−y)=2(0.3165−1.0)=−1.3670. Since y^=z2 exactly, ∂L/∂z2=−1.3670 too. The sign is negative because the prediction is too low: increasing it would reduce the loss.
Second layer parameters. Since z2=W2a1+b2, we have ∂z2/∂W2=a1 and ∂z2/∂b2=1. Multiply through:
Read that: the gradient for a weight is the incoming error signal times the activation that weight multiplied. A weight attached to a large activation gets a large update; a weight attached to a near-zero activation barely moves. That is the general pattern for every weight in every layer.
Push the signal back through the second layer. ∂z2/∂a1=W2=0.8, so
Through the sigmoid. The sigmoid's derivative is a1(1−a1)=0.6457×0.3543=0.2288, so
First layer parameters. ∂z1/∂W1=x=1.0 and ∂z1/∂b1=1:
All four gradients, from one forward pass and one backward pass. Update each parameter by subtracting a small multiple of its gradient and the loss goes down.
Two things this worked example makes visible
The forward values are needed again. Computing ∂L/∂W2 required a1, and the sigmoid's local derivative required a1 again. This is why frameworks cache activations during the forward pass and why memory use scales with depth and batch size. Discard them and you would have to recompute the forward pass at every backward step.
The signal shrinks as it travels. It arrived at the output as −1.3670 and reached the first layer as −0.2502 — smaller by a factor of 5.5, after crossing a single sigmoid. The culprit is that σ′ can never exceed 0.25 and is usually much less. Cross ten sigmoids and the signal is multiplied by ten such factors: 0.2510≈10−6 in the best case. The early layers receive essentially nothing and stop learning. That is the vanishing gradient problem, and it is visible right here in the arithmetic, not in some abstract theory.
The reverse can also happen. If the weights are large, each backward step multiplies by a large number, and the product explodes to infinity within a few layers. Exploding gradients produce a loss that jumps to nan in a single step.
| Symptom | What the chain rule is doing | Usual fix |
|---|---|---|
| Early layers barely change | Many local derivatives below 1, multiplied together | ReLU instead of sigmoid, residual connections, normalisation layers |
Loss becomes nan | Many local derivatives above 1, multiplied together | Gradient clipping, smaller initialisation, lower learning rate |
| Whole units stop updating | ReLU stuck at negative input, local derivative exactly 0 | Leaky ReLU, lower learning rate, better initialisation |
Doing it in code, and proving it is right
1import numpy as np23def sigmoid(z):4 return 1.0 / (1.0 + np.exp(-z))56def forward(params, x, y):7 W1, b1, W2, b2 = params8 z1 = W1 * x + b19 a1 = sigmoid(z1)10 z2 = W2 * a1 + b211 loss = (z2 - y) ** 212 cache = (z1, a1, z2)13 return loss, cache1415def backward(params, cache, x, y):16 W1, b1, W2, b2 = params17 z1, a1, z2 = cache1819 dL_dz2 = 2.0 * (z2 - y) # -1.366920 dL_dW2 = dL_dz2 * a1 # -0.882621 dL_db2 = dL_dz2 # -1.36692223 dL_da1 = dL_dz2 * W2 # -1.093624 dL_dz1 = dL_da1 * a1 * (1 - a1) # -0.250225 dL_dW1 = dL_dz1 * x # -0.250226 dL_db1 = dL_dz1 # -0.25022728 return np.array([dL_dW1, dL_db1, dL_dW2, dL_db2])2930params = np.array([0.5, 0.1, 0.8, -0.2])31x, y = 1.0, 1.03233loss, cache = forward(params, x, y)34analytic = backward(params, cache, x, y)3536# Gradient check: compare against central differences.37numeric = np.zeros_like(params)38h = 1e-639for i in range(len(params)):40 up, down = params.copy(), params.copy()41 up[i] += h42 down[i] -= h43 numeric[i] = (forward(up, x, y)[0] - forward(down, x, y)[0]) / (2 * h)4445print(analytic)46print(numeric)47print("max abs difference:", np.max(np.abs(analytic - numeric)))The comments show what the code prints. Two of them differ from the hand calculation in the fourth decimal place (−1.3669 against −1.3670, −0.8826 against −0.8827) only because the hand calculation rounded each intermediate value as it went.
That gradient check is worth building into any hand-written backward pass. The two arrays should agree to roughly six decimal places. If they do not, the analytic derivative has a bug — and a wrong gradient does not raise an exception, it just trains a worse model, slowly, with no indication of why. Run the check once on a tiny input, then delete it or hide it behind a flag, because it is far too slow to run during real training.
What this changes about how you read a model
Once you see a network as a pipeline of simple operations, the composite expression at the top of this page stops being intimidating. You never differentiate it as a whole. You differentiate each step locally — a multiplication contributes the other factor, an addition contributes 1, a sigmoid contributes a(1−a) — and the chain rule multiplies those local pieces together on the way back.
That is also why automatic differentiation works at all. A framework does not do algebra on your model. It records the sequence of elementary operations as you run the forward pass, then walks that record backwards applying the local derivative of each operation. Add a new layer type and, as long as you can state its local derivative, everything else keeps working unchanged.
The practical payoff is diagnostic. When training misbehaves, print the gradient magnitudes layer by layer. If they shrink by an order of magnitude per layer going backwards, you have a saturation problem and the fix is architectural. If they grow, you need clipping or a smaller initialisation. If they are exactly zero somewhere, a unit has died. All three readings come straight from understanding that the gradient at layer k is a product of every local derivative between layer k and the loss.