Mathematics for Machine Learning

Functions, Derivatives, and Gradients


Imagine tuning a single parameter of a model by hand. You set the weight to 0.5, run the training set through, and the loss comes back as 4.21. You try 0.6 — loss goes up to 4.88. So you go the other way: 0.4 gives 3.71. Better. Try 0.3, get 3.44. Try 0.2, get 3.39. Try 0.1, and now it is 3.56, worse again. So the best value is somewhere near 0.2, and you start bisecting.

That took six full passes over the data to locate one parameter approximately. A small neural network has ten thousand parameters. At six passes each, you would need sixty thousand passes to do one round of tuning, and the parameters interact, so you would have to do it again from scratch. This approach does not scale to two parameters, let alone a million.

The problem is that each trial gives you one bit of information — better or worse — and you paid for it with a full pass over the data. What you actually want, standing at weight 0.5, is a number that tells you which way to move and how urgently, computed from that one position. That number is the derivative, and computing it costs roughly the same as evaluating the loss once.

Probing the loss by hand, one weight at a time3.563.393.443.714.214.88012345lowest, w = 0.2youstarted hereWeights 0.1 to 0.6 left to right; the sign flips between 0.2 and 0.3.
The derivative is what you were estimating by hand — its sign says which way to move, its size says how far.

The derivative is a rate of change you can measure at a point

Start with the honest, clumsy version. Pick a function, pick a point, nudge the input by a small amount hh, and see how much the output moved per unit of nudge:

f(x+h)−f(x)h\frac{f(x+h) - f(x)}{h}

This is the slope of the straight line through two nearby points on the curve — the average rate of change over that little interval. It depends on hh, which is unsatisfying. So shrink hh towards zero and see whether the answer settles down:

f′(x)=lim⁡h→0f(x+h)−f(x)hf'(x) = \lim_{h \to 0} \frac{f(x+h) - f(x)}{h}

Watch it settle for f(x)=x2f(x) = x^2 at x=3x = 3:

hhf(3+h)f(3+h)f(3+h)−9h\dfrac{f(3+h) - 9}{h}
1167
0.19.616.1
0.019.06016.01
0.0019.0060016.001

It is converging to 6. And the general rule for x2x^2 gives f′(x)=2xf'(x) = 2x, which at x=3x = 3 is exactly 6.

The meaning: at x=3x = 3, this function is climbing at six units of output per unit of input. If you nudge xx up by 0.01, expect the output to rise by about 6×0.01=0.066 \times 0.01 = 0.06. That is a prediction you can make from one evaluation instead of a search.

The derivative is the exchange rate between a small change in the input and the resulting change in the output.

Two readings of the same number are both useful. Geometrically, f′(x)f'(x) is the slope of the line that just touches the curve at xx — the tangent. Practically, it is a sensitivity: how much does this output care about this input, right here?

The rules you will actually use

RuleStatementWorked instance
Powerddxxn=nxn−1\frac{d}{dx}x^n = nx^{n-1}ddxx5=5x4\frac{d}{dx}x^5 = 5x^4
Constant multipleddx[cf]=cf′\frac{d}{dx}[cf] = cf'ddx7x3=21x2\frac{d}{dx}7x^3 = 21x^2
Sumddx[f+g]=f′+g′\frac{d}{dx}[f+g] = f' + g'ddx(x2+3x)=2x+3\frac{d}{dx}(x^2 + 3x) = 2x + 3
Productddx[fg]=f′g+fg′\frac{d}{dx}[fg] = f'g + fg'ddx(x2ex)=2xex+x2ex\frac{d}{dx}(x^2 e^x) = 2xe^x + x^2e^x
Quotientddx[fg]=f′g−fg′g2\frac{d}{dx}\left[\frac{f}{g}\right] = \frac{f'g - fg'}{g^2}ddxxx+1=1(x+1)2\frac{d}{dx}\frac{x}{x+1} = \frac{1}{(x+1)^2}
Chainddxf(g(x))=f′(g(x)) g′(x)\frac{d}{dx}f(g(x)) = f'(g(x))\,g'(x)ddx(3x+1)4=12(3x+1)3\frac{d}{dx}(3x+1)^4 = 12(3x+1)^3

The chain rule is the one that carries all the weight in machine learning, because a model is nothing but functions stacked on functions. Its statement is: to differentiate a composition, differentiate the outer function while treating the inner one as a single blob, then multiply by the derivative of the inner one. In (3x+1)4(3x+1)^4, the outer operation is "raise to the fourth", giving 4(3x+1)34(3x+1)^3, and the inner is 3x+13x+1, whose derivative is 3. Multiply: 12(3x+1)312(3x+1)^3.

The specific functions that appear in models

FunctionDerivativeWhy it matters
exe^xexe^xIts own derivative — why exponentials are everywhere in the algebra
ln⁡x\ln x1/x1/xTurns products into sums; the basis of log-likelihood losses
σ(x)=11+e−x\sigma(x) = \frac{1}{1+e^{-x}}σ(x)(1−σ(x))\sigma(x)(1 - \sigma(x))Squashes to (0,1)(0,1) for probabilities
tanh⁡x\tanh x1−tanh⁡2x1 - \tanh^2 xZero-centred alternative to sigmoid
ReLU(x)=max⁡(0,x)\mathrm{ReLU}(x) = \max(0,x)1 if x>0x>0, 0 if x<0x<0Constant gradient, so no squashing

The sigmoid derivative deserves a moment because it explains a real training failure. Since σ\sigma always lies between 0 and 1, the product σ(1−σ)\sigma(1-\sigma) is largest at σ=0.5\sigma = 0.5, where it equals 0.250.25, and it collapses towards zero at both extremes. At x=5x = 5, σ(5)≈0.9933\sigma(5) \approx 0.9933 and the derivative is 0.9933×0.0067≈0.00660.9933 \times 0.0067 \approx 0.0066. So a neuron that has become confident has almost no gradient: the signal telling it how to improve has been multiplied by 0.0066. Stack a few of those and the signal reaching the early layers is effectively zero. That is the vanishing gradient problem, and it is visible directly in this one formula.

ReLU has the opposite property — its derivative is exactly 1 for any positive input, so gradients pass through undiminished. It has no derivative at x=0x = 0, because the slope jumps from 0 to 1 there. Frameworks resolve this by convention, returning 0 (sometimes 1) at exactly zero. It is mathematically a fudge and practically irrelevant: the probability of landing exactly on 0 in floating-point arithmetic is negligible, and the function is continuous, so nothing blows up.

More than one input: partial derivatives

Real loss functions do not take one number. They take every parameter in the model. So "the slope" is not well defined any more — slope in which direction?

The fix is to take one variable at a time. A partial derivative asks: if I change this one input and freeze every other input, how does the output respond? The notation switches from dd to ∂\partial to signal that the other variables are being held still.

Take f(x,y)=x2y+3y2f(x, y) = x^2y + 3y^2.

For ∂f/∂x\partial f/\partial x, treat yy as a constant. Then x2yx^2 y differentiates to 2xy2xy, and 3y23y^2 is a constant, so it differentiates to 0:

∂f∂x=2xy\frac{\partial f}{\partial x} = 2xy

For ∂f/∂y\partial f/\partial y, treat xx as a constant. Now x2yx^2y is a constant times yy, giving x2x^2, and 3y23y^2 gives 6y6y:

∂f∂y=x2+6y\frac{\partial f}{\partial y} = x^2 + 6y

Evaluate both at the point (1,2)(1, 2): ∂f/∂x=2(1)(2)=4\partial f/\partial x = 2(1)(2) = 4, and ∂f/∂y=1+12=13\partial f/\partial y = 1 + 12 = 13. So at that point, the function is more than three times as sensitive to yy as to xx. Nudging yy buys you much more change than nudging xx.

The gradient collects them all

∇f=[∂f∂x1∂f∂x2⋮∂f∂xn]\nabla f = \begin{bmatrix} \dfrac{\partial f}{\partial x_1} \\[6pt] \dfrac{\partial f}{\partial x_2} \\[4pt] \vdots \\[2pt] \dfrac{\partial f}{\partial x_n} \end{bmatrix}

The gradient is just the list of partial derivatives arranged as a vector. At (1,2)(1,2) our example has ∇f=[4,13]T\nabla f = [4, 13]^T. But a vector has a direction, and this one has a remarkable property:

The gradient points in the direction of steepest ascent, and its length is how steep that ascent is.

Both halves matter. The direction tells you where to go to increase the function fastest — so its negative, −∇f-\nabla f, is the fastest way down, which is exactly what you want when the function is a loss. The magnitude ∥∇f∥\|\nabla f\| tells you how urgent the situation is: a long gradient means you are on a steep slope far from anywhere flat, and a near-zero gradient means the surface has levelled off.

For our example, ∥∇f∥=42+132=185≈13.60\|\nabla f\| = \sqrt{4^2 + 13^2} = \sqrt{185} \approx 13.60. That is the maximum rate of increase available at that point in any direction.

Rate of change in a direction you choose

Suppose you do not want the steepest direction — you want to know how fast ff changes if you move along some specific unit vector u\mathbf{u}. The answer is a dot product:

Duf=∇f⋅uD_{\mathbf{u}}f = \nabla f \cdot \mathbf{u}

Using ∇f=[4,13]T\nabla f = [4, 13]^T and the unit direction u=[0.6,0.8]T\mathbf{u} = [0.6, 0.8]^T (its length is 0.36+0.64=1\sqrt{0.36 + 0.64} = 1, as required):

Duf=4(0.6)+13(0.8)=2.4+10.4=12.8D_{\mathbf{u}}f = 4(0.6) + 13(0.8) = 2.4 + 10.4 = 12.8

Moving along u\mathbf{u}, the function climbs at 12.8 per unit distance — slightly less than the 13.60 you would get by going in the gradient's own direction, because u\mathbf{u} is not quite aligned with it. This also explains why the gradient is the steepest direction: the dot product ∇f⋅u=∥∇f∥cos⁡θ\nabla f \cdot \mathbf{u} = \|\nabla f\|\cos\theta is largest when cos⁡θ=1\cos\theta = 1, that is, when u\mathbf{u} points the same way as ∇f\nabla f. And when u\mathbf{u} is perpendicular to the gradient, the rate is zero — you are walking along a contour line, staying at the same height.

The unit-length requirement on u\mathbf{u} is not decoration. If you pass a vector of length 3, you get three times the answer, because you have measured the change over three units of distance rather than one.

Where the gradient goes quiet: critical points

A critical point is where ∇f=0\nabla f = \mathbf{0} — every partial derivative vanishes, so no direction offers first-order improvement. Every minimum, maximum and saddle point is a critical point, which is why optimisation is fundamentally a hunt for places where the gradient is zero.

Take f(x,y)=x2+y2−4x−6y+13f(x, y) = x^2 + y^2 - 4x - 6y + 13. Set both partials to zero:

∂f∂x=2x−4=0  ⇒  x=2,∂f∂y=2y−6=0  ⇒  y=3\frac{\partial f}{\partial x} = 2x - 4 = 0 \;\Rightarrow\; x = 2, \qquad \frac{\partial f}{\partial y} = 2y - 6 = 0 \;\Rightarrow\; y = 3

The only critical point is (2,3)(2, 3), where f=4+9−8−18+13=0f = 4 + 9 - 8 - 18 + 13 = 0.

A zero gradient does not by itself say what kind of point you found. In one dimension the second derivative settles it: f′′(x)>0f''(x) > 0 means the curve is bending upwards, so you are in a valley; f′′(x)<0f''(x) < 0 means a peak. For f(x)=x3−3xf(x) = x^3 - 3x, setting f′(x)=3x2−3=0f'(x) = 3x^2 - 3 = 0 gives x=±1x = \pm 1, and f′′(x)=6xf''(x) = 6x is +6+6 at x=1x = 1 (a minimum) and −6-6 at x=−1x = -1 (a maximum).

In higher dimensions you need the whole matrix of second derivatives, and its eigenvalue signs decide the answer: all positive is a minimum, all negative a maximum, mixed signs a saddle — a point that is a minimum along some directions and a maximum along others. In a model with millions of parameters, saddles vastly outnumber true minima, because getting every single direction to curve the same way becomes wildly unlikely as dimensions grow. This is a genuinely different picture from the "get stuck in a local minimum" story people tend to carry over from one-dimensional intuition.

Getting derivatives out of a computer

Python
import numpy as npdef f(v):    x, y = v    return x**2 * y + 3 * y**2def grad_analytic(v):    x, y = v    return np.array([2 * x * y, x**2 + 6 * y])def grad_numeric(f, v, h=1e-5):    """Central differences: more accurate than the forward version,    with error O(h^2) instead of O(h)."""    g = np.zeros_like(v, dtype=float)    for i in range(len(v)):        step = np.zeros_like(v, dtype=float)        step[i] = h        g[i] = (f(v + step) - f(v - step)) / (2 * h)    return gp = np.array([1.0, 2.0])print(grad_analytic(p))            # [ 4. 13.]print(grad_numeric(f, p))          # [ 4. 13.] to about 9 decimal places

The numeric version is not how models are trained — it needs two function evaluations per parameter, so a million parameters means two million forward passes. It is, however, exactly how you check that a hand-written analytic gradient is correct, and that check has saved more debugging hours than almost any other habit. Compare the two and expect agreement to five or six decimal places; anything worse means the analytic derivative has a bug.

Note the choice of hh. Too large and the finite difference is a poor approximation to the limit. Too small and subtracting two nearly-equal floating-point numbers destroys precision. Around 10−510^{-5} for central differences is the usual sweet spot.

Production frameworks use automatic differentiation instead: they record every elementary operation during the forward computation and then apply the chain rule mechanically backwards through that record. The result is exact to machine precision, and it costs roughly the same as one extra forward pass regardless of how many parameters there are.

The mistakes that cost people days

MistakeWhat goes wrong
Adding the gradient instead of subtracting itThe loss climbs steadily instead of falling. The gradient points uphill; you must move against it.
Treating ∇f\nabla f as a numberIt is a vector with one entry per parameter. Shape errors, or worse, silent broadcasting.
Differentiating with respect to the dataThe data is fixed; the parameters are what you can change. The gradient you want is ∂L/∂θ\partial L/\partial \theta, never ∂L/∂x\partial L/\partial x (except in adversarial-example work, where that is precisely the point).
Forgetting the inner derivative in a chainGradients come out scaled by a constant factor. Training still sort of works, which makes it hard to spot.
Reading a zero gradient as "done"It could be a saddle, or a flat plateau, or a dead ReLU that will never recover.
Not checking against finite differencesA wrong analytic gradient produces a model that trains slowly and badly, with no error message.

What this buys you in practice

Go back to the hand-tuning at the start. With a derivative, standing at weight 0.5 with a loss of 4.21, you compute dL/dwdL/dw once and it comes back as, say, +2.1+2.1. Positive means increasing ww increases the loss, so you decrease ww. The magnitude 2.1 tells you the slope is moderate, so a step proportional to it is reasonable. One evaluation, and you know both the direction and the scale. Do that for all ten thousand parameters simultaneously — which is exactly what one backward pass gives you — and the whole model improves at once.

The deeper habit is to stop thinking of a derivative as a symbolic exercise and start reading it as a sensitivity report. When a model will not train, the useful question is nearly always about gradients: are they zero, in which case something has saturated or died; are they enormous, in which case the scale of the inputs or the initialisation is wrong; or are they fine and the step size is wrong. All three diagnoses come from looking at ∇f\nabla f as a set of numbers with meanings, rather than as a formula you derived once and stopped thinking about.