Course Content
Mathematics for Machine Learning
5 sections · 13 lessons
Functions, Derivatives, and Gradients
Imagine tuning a single parameter of a model by hand. You set the weight to 0.5, run the training set through, and the loss comes back as 4.21. You try 0.6 — loss goes up to 4.88. So you go the other way: 0.4 gives 3.71. Better. Try 0.3, get 3.44. Try 0.2, get 3.39. Try 0.1, and now it is 3.56, worse again. So the best value is somewhere near 0.2, and you start bisecting.
That took six full passes over the data to locate one parameter approximately. A small neural network has ten thousand parameters. At six passes each, you would need sixty thousand passes to do one round of tuning, and the parameters interact, so you would have to do it again from scratch. This approach does not scale to two parameters, let alone a million.
The problem is that each trial gives you one bit of information — better or worse — and you paid for it with a full pass over the data. What you actually want, standing at weight 0.5, is a number that tells you which way to move and how urgently, computed from that one position. That number is the derivative, and computing it costs roughly the same as evaluating the loss once.
The derivative is a rate of change you can measure at a point
Start with the honest, clumsy version. Pick a function, pick a point, nudge the input by a small amount h, and see how much the output moved per unit of nudge:
This is the slope of the straight line through two nearby points on the curve — the average rate of change over that little interval. It depends on h, which is unsatisfying. So shrink h towards zero and see whether the answer settles down:
Watch it settle for f(x)=x2 at x=3:
| h | f(3+h) | hf(3+h)−9 |
|---|---|---|
| 1 | 16 | 7 |
| 0.1 | 9.61 | 6.1 |
| 0.01 | 9.0601 | 6.01 |
| 0.001 | 9.006001 | 6.001 |
It is converging to 6. And the general rule for x2 gives f′(x)=2x, which at x=3 is exactly 6.
The meaning: at x=3, this function is climbing at six units of output per unit of input. If you nudge x up by 0.01, expect the output to rise by about 6×0.01=0.06. That is a prediction you can make from one evaluation instead of a search.
The derivative is the exchange rate between a small change in the input and the resulting change in the output.
Two readings of the same number are both useful. Geometrically, f′(x) is the slope of the line that just touches the curve at x — the tangent. Practically, it is a sensitivity: how much does this output care about this input, right here?
The rules you will actually use
| Rule | Statement | Worked instance |
|---|---|---|
| Power | dxdxn=nxn−1 | dxdx5=5x4 |
| Constant multiple | dxd[cf]=cf′ | dxd7x3=21x2 |
| Sum | dxd[f+g]=f′+g′ | dxd(x2+3x)=2x+3 |
| Product | dxd[fg]=f′g+fg′ | dxd(x2ex)=2xex+x2ex |
| Quotient | dxd[gf]=g2f′g−fg′ | dxdx+1x=(x+1)21 |
| Chain | dxdf(g(x))=f′(g(x))g′(x) | dxd(3x+1)4=12(3x+1)3 |
The chain rule is the one that carries all the weight in machine learning, because a model is nothing but functions stacked on functions. Its statement is: to differentiate a composition, differentiate the outer function while treating the inner one as a single blob, then multiply by the derivative of the inner one. In (3x+1)4, the outer operation is "raise to the fourth", giving 4(3x+1)3, and the inner is 3x+1, whose derivative is 3. Multiply: 12(3x+1)3.
The specific functions that appear in models
| Function | Derivative | Why it matters |
|---|---|---|
| ex | ex | Its own derivative — why exponentials are everywhere in the algebra |
| lnx | 1/x | Turns products into sums; the basis of log-likelihood losses |
| σ(x)=1+e−x1 | σ(x)(1−σ(x)) | Squashes to (0,1) for probabilities |
| tanhx | 1−tanh2x | Zero-centred alternative to sigmoid |
| ReLU(x)=max(0,x) | 1 if x>0, 0 if x<0 | Constant gradient, so no squashing |
The sigmoid derivative deserves a moment because it explains a real training failure. Since σ always lies between 0 and 1, the product σ(1−σ) is largest at σ=0.5, where it equals 0.25, and it collapses towards zero at both extremes. At x=5, σ(5)≈0.9933 and the derivative is 0.9933×0.0067≈0.0066. So a neuron that has become confident has almost no gradient: the signal telling it how to improve has been multiplied by 0.0066. Stack a few of those and the signal reaching the early layers is effectively zero. That is the vanishing gradient problem, and it is visible directly in this one formula.
ReLU has the opposite property — its derivative is exactly 1 for any positive input, so gradients pass through undiminished. It has no derivative at x=0, because the slope jumps from 0 to 1 there. Frameworks resolve this by convention, returning 0 (sometimes 1) at exactly zero. It is mathematically a fudge and practically irrelevant: the probability of landing exactly on 0 in floating-point arithmetic is negligible, and the function is continuous, so nothing blows up.
More than one input: partial derivatives
Real loss functions do not take one number. They take every parameter in the model. So "the slope" is not well defined any more — slope in which direction?
The fix is to take one variable at a time. A partial derivative asks: if I change this one input and freeze every other input, how does the output respond? The notation switches from d to ∂ to signal that the other variables are being held still.
Take f(x,y)=x2y+3y2.
For ∂f/∂x, treat y as a constant. Then x2y differentiates to 2xy, and 3y2 is a constant, so it differentiates to 0:
For ∂f/∂y, treat x as a constant. Now x2y is a constant times y, giving x2, and 3y2 gives 6y:
Evaluate both at the point (1,2): ∂f/∂x=2(1)(2)=4, and ∂f/∂y=1+12=13. So at that point, the function is more than three times as sensitive to y as to x. Nudging y buys you much more change than nudging x.
The gradient collects them all
The gradient is just the list of partial derivatives arranged as a vector. At (1,2) our example has ∇f=[4,13]T. But a vector has a direction, and this one has a remarkable property:
The gradient points in the direction of steepest ascent, and its length is how steep that ascent is.
Both halves matter. The direction tells you where to go to increase the function fastest — so its negative, −∇f, is the fastest way down, which is exactly what you want when the function is a loss. The magnitude ∥∇f∥ tells you how urgent the situation is: a long gradient means you are on a steep slope far from anywhere flat, and a near-zero gradient means the surface has levelled off.
For our example, ∥∇f∥=42+132=185≈13.60. That is the maximum rate of increase available at that point in any direction.
Rate of change in a direction you choose
Suppose you do not want the steepest direction — you want to know how fast f changes if you move along some specific unit vector u. The answer is a dot product:
Using ∇f=[4,13]T and the unit direction u=[0.6,0.8]T (its length is 0.36+0.64=1, as required):
Moving along u, the function climbs at 12.8 per unit distance — slightly less than the 13.60 you would get by going in the gradient's own direction, because u is not quite aligned with it. This also explains why the gradient is the steepest direction: the dot product ∇f⋅u=∥∇f∥cosθ is largest when cosθ=1, that is, when u points the same way as ∇f. And when u is perpendicular to the gradient, the rate is zero — you are walking along a contour line, staying at the same height.
The unit-length requirement on u is not decoration. If you pass a vector of length 3, you get three times the answer, because you have measured the change over three units of distance rather than one.
Where the gradient goes quiet: critical points
A critical point is where ∇f=0 — every partial derivative vanishes, so no direction offers first-order improvement. Every minimum, maximum and saddle point is a critical point, which is why optimisation is fundamentally a hunt for places where the gradient is zero.
Take f(x,y)=x2+y2−4x−6y+13. Set both partials to zero:
The only critical point is (2,3), where f=4+9−8−18+13=0.
A zero gradient does not by itself say what kind of point you found. In one dimension the second derivative settles it: f′′(x)>0 means the curve is bending upwards, so you are in a valley; f′′(x)<0 means a peak. For f(x)=x3−3x, setting f′(x)=3x2−3=0 gives x=±1, and f′′(x)=6x is +6 at x=1 (a minimum) and −6 at x=−1 (a maximum).
In higher dimensions you need the whole matrix of second derivatives, and its eigenvalue signs decide the answer: all positive is a minimum, all negative a maximum, mixed signs a saddle — a point that is a minimum along some directions and a maximum along others. In a model with millions of parameters, saddles vastly outnumber true minima, because getting every single direction to curve the same way becomes wildly unlikely as dimensions grow. This is a genuinely different picture from the "get stuck in a local minimum" story people tend to carry over from one-dimensional intuition.
Getting derivatives out of a computer
1import numpy as np23def f(v):4 x, y = v5 return x**2 * y + 3 * y**267def grad_analytic(v):8 x, y = v9 return np.array([2 * x * y, x**2 + 6 * y])1011def grad_numeric(f, v, h=1e-5):12 """Central differences: more accurate than the forward version,13 with error O(h^2) instead of O(h)."""14 g = np.zeros_like(v, dtype=float)15 for i in range(len(v)):16 step = np.zeros_like(v, dtype=float)17 step[i] = h18 g[i] = (f(v + step) - f(v - step)) / (2 * h)19 return g2021p = np.array([1.0, 2.0])22print(grad_analytic(p)) # [ 4. 13.]23print(grad_numeric(f, p)) # [ 4. 13.] to about 9 decimal placesThe numeric version is not how models are trained — it needs two function evaluations per parameter, so a million parameters means two million forward passes. It is, however, exactly how you check that a hand-written analytic gradient is correct, and that check has saved more debugging hours than almost any other habit. Compare the two and expect agreement to five or six decimal places; anything worse means the analytic derivative has a bug.
Note the choice of h. Too large and the finite difference is a poor approximation to the limit. Too small and subtracting two nearly-equal floating-point numbers destroys precision. Around 10−5 for central differences is the usual sweet spot.
Production frameworks use automatic differentiation instead: they record every elementary operation during the forward computation and then apply the chain rule mechanically backwards through that record. The result is exact to machine precision, and it costs roughly the same as one extra forward pass regardless of how many parameters there are.
The mistakes that cost people days
| Mistake | What goes wrong |
|---|---|
| Adding the gradient instead of subtracting it | The loss climbs steadily instead of falling. The gradient points uphill; you must move against it. |
| Treating ∇f as a number | It is a vector with one entry per parameter. Shape errors, or worse, silent broadcasting. |
| Differentiating with respect to the data | The data is fixed; the parameters are what you can change. The gradient you want is ∂L/∂θ, never ∂L/∂x (except in adversarial-example work, where that is precisely the point). |
| Forgetting the inner derivative in a chain | Gradients come out scaled by a constant factor. Training still sort of works, which makes it hard to spot. |
| Reading a zero gradient as "done" | It could be a saddle, or a flat plateau, or a dead ReLU that will never recover. |
| Not checking against finite differences | A wrong analytic gradient produces a model that trains slowly and badly, with no error message. |
What this buys you in practice
Go back to the hand-tuning at the start. With a derivative, standing at weight 0.5 with a loss of 4.21, you compute dL/dw once and it comes back as, say, +2.1. Positive means increasing w increases the loss, so you decrease w. The magnitude 2.1 tells you the slope is moderate, so a step proportional to it is reasonable. One evaluation, and you know both the direction and the scale. Do that for all ten thousand parameters simultaneously — which is exactly what one backward pass gives you — and the whole model improves at once.
The deeper habit is to stop thinking of a derivative as a symbolic exercise and start reading it as a sensitivity report. When a model will not train, the useful question is nearly always about gradients: are they zero, in which case something has saturated or died; are they enormous, in which case the scale of the inputs or the initialisation is wrong; or are they fine and the step size is wrong. All three diagnoses come from looking at ∇f as a set of numbers with meanings, rather than as a formula you derived once and stopped thinking about.