Course Content
Mathematics for Machine Learning
5 sections · 13 lessons
Loss Functions — Teaching a Model What "Wrong" Means
Two models predict house prices. Model A is off by £5,000 on every single house. Model B is exactly right on 90 houses out of 100 and off by £200,000 on the other ten. Which model is better?
You cannot answer that. Not because you lack information — you have every prediction and every true price — but because "better" is not a property of the numbers. It is a decision you have to make and write down. A mortgage lender pricing risk and an estate agent writing brochures would answer differently, and both would be right.
Try the obvious rule: count how many predictions are exactly correct. Model B wins, 90 to nothing. But that rule is useless for training. With continuous outputs you almost never hit a value exactly, so the count is zero for nearly every model you could build, and being off by £1 scores the same as being off by £1 million. Worse, the count is a staircase: nudge a weight by a hair and it does not move. Gradient-based learning needs a slope to slide down, and a staircase has none.
So you need something else: a function that eats one true value and one prediction and returns a single number saying how bad that was, smoothly enough to have a slope. That is a loss function. And here is what catches people out: the model does not care what you meant. It will push that number down with total, literal obedience.
Two levels: one example, and the whole dataset
The per-example loss ℓ(y,y^) scores a single prediction: non-negative, zero when perfect, larger as the prediction gets worse. It knows nothing about the model or the rest of the data — just one truth and one guess.
The empirical risk is that averaged over your training set:
Here θ stands for every parameter you are allowed to change, and training means finding the θ that makes L(θ) smallest. The word empirical matters: this is the loss on the finite pile of data you happen to have, standing in for the loss you actually care about, the average over all the data you will never see. That gap is why a model can look brilliant in training and fall over in production.
Dividing by n is not cosmetic. Without it, doubling your dataset doubles every gradient, so a learning rate that worked on 10,000 rows explodes on 20,000.
A model does not learn the task you had in mind. It learns to make one number small. Everything else is your responsibility.
What makes a loss usable
- Differentiable almost everywhere. Isolated kinks are fine — like ∣x∣ at zero — flat regions are not. Accuracy fails badly: it is a step function whose derivative is zero wherever it exists, so the optimiser is told "moving does nothing" and stands still.
- Aligned with the goal. The loss is the only channel through which your intentions reach the model. If a missed fraud costs £10,000 and a false alarm costs a phone call, and the loss treats both identically, the model will happily trade fraud for phone calls.
- Well-scaled. Loss values around 106 produce enormous gradients and demand an absurdly small learning rate; values around 10−8 vanish into floating-point noise. Predicting prices in pounds rather than thousands multiplies squared-error loss values by a million and their gradients by a thousand.
- Convex in the prediction, ideally. Convex means bowl-shaped: one bottom, no local traps. A deep network destroys global convexity anyway, but a convex per-prediction loss guarantees the loss is not adding minima of its own on top of the network's.
Regression losses: the choice is really about outliers
Mean squared error
Square each error so overshooting by 3 and undershooting by 3 cost the same, then average. Read the gradient in words: the push each example applies is proportional to how wrong it is. An example off by 10 pulls ten times harder than one off by 1. On clean data that is reasonable. On dirty data it is a trap.
The concrete cost of squaring
Take ten predictions. Nine are off by exactly 1. One — a mistyped data entry, say — is off by 10.
MSE: the nine contribute 9×12=9, the bad one contributes 102=100. Total 109, so MSE=10.9. The single outlier is 91.7% of the entire loss.
MAE (mean absolute error, n1∑∣yi−y^i∣): the nine contribute 9, the outlier contributes 10, so MAE=1.9. The outlier is 52.6% of the loss.
Same data, same predictions. Under MSE nine-tenths of the loss comes from one possibly-corrupt row, and because the MSE gradient grows with the error, that row also supplies more than half of the push on the model (10 of the 19 units of error). The model bends itself out of shape to accommodate it. Under MAE the row counts as ten ordinary rows in the loss, but its push is the same ±1 as every other row's — heavy, but not dictatorial.
Mean absolute error and its awkward corner
MAE's gradient is ±1 regardless of the size of the error: only the direction changes. That is what makes it robust and also what makes it awkward. As you converge the gradient never shrinks, so the optimiser keeps taking full-size steps and bounces around the minimum instead of settling into it; you need a decaying learning rate to compensate.
MAE is also not differentiable at zero, the point where the prediction is exactly right. Formally you use a subgradient — any slope between −1 and +1 is valid there. Practically every framework returns 0 and nothing bad happens, because landing on an exact tie is vanishingly rare in floating point.
Huber loss: the practical compromise
Writing the error as a=y−y^:
Quadratic for small errors, linear for large ones. The −21δ is not decoration: it makes the two pieces meet at ∣a∣=δ with the same value and the same slope, so there is no jump and no kink. The gradient rises with the error up to a magnitude of δ, then stays capped.
So Huber behaves like MSE near the answer, where you want smooth convergence, and like MAE out in the tails, where you want protection. The parameter δ is the line you draw between "ordinary error" and "outlier"; a common starting point is the 90th percentile of absolute residuals from a quick first fit. Large δ recovers MSE, small δ a scaled MAE.
| Loss | Per-example form | Gradient behaviour | Outlier sensitivity | Differentiable | Pulls prediction towards |
|---|---|---|---|---|---|
| MSE | (y−y^)2 | Grows with the error, unbounded | Very high — errors are squared | Everywhere | the mean |
| MAE | ∣y−y^∣ | Constant ±1, never shrinks | Low — errors counted once | Everywhere except y^=y | the median |
| Huber(δ) | quadratic inside δ, linear outside | Grows to ±δ, then capped | Low, and smooth near zero | Everywhere, including the joins | between the two |
Why the last column is the real reason to choose
Ask what single constant c minimises each loss on a fixed set of numbers. Differentiate ∑i(yi−c)2 and set to zero: −2∑i(yi−c)=0, giving c=n1∑yi, the mean. The derivative of ∑i∣yi−c∣ counts how many points sit above c minus how many sit below; that is zero when the counts match, which is the median.
On the data 2,4,5,6,100: MSE drives your prediction towards 23.4, MAE towards 5. Neither is wrong. Forecasting warehouse stock you will later sum? You need unbiased totals, so you need the mean, so you need MSE. Telling a customer how long delivery takes and wanting to be right half the time? That is the median, so MAE.
Choosing between squared and absolute error is not a technicality about outliers. It decides whether your model reports the mean or the median.
Classification: why squared error is the wrong tool
The obvious move is to encode the label as 0 or 1, output a probability p^, and use (y−p^)2. It fails for two reasons.
A classifier produces a raw score z (a logit) and squashes it with the sigmoid p^=1/(1+e−z), whose derivative is p^(1−p^). By the chain rule, ∂z∂(p^−y)2=2(p^−y)p^(1−p^). Suppose the true label is 1 and the model confidently says p^=0.01 — as wrong as it gets. Then p^(1−p^)=0.0099 and the gradient is 2×(−0.99)×0.0099≈−0.0196. Tiny. The model is maximally wrong and gets almost no correction, because the sigmoid has saturated: learning stalls exactly where it is most needed. Second, squared error composed with a sigmoid is not convex even for a single-layer model, so you inherit local minima you did not need.
Binary cross-entropy, and where it comes from
This is not arbitrary; it falls out of maximum likelihood. Treat each label as a coin flip whose bias the model predicts. The probability the model assigns to the label actually observed is p^y(1−p^)1−y — check it: y=1 gives p^, y=0 gives 1−p^. Assume independence, so the dataset's probability is the product of those terms. Products of thousands of numbers below 1 underflow to zero, so take the logarithm, turning the product into a sum; maximising is minimising the negative, so flip the sign; divide by n. That is exactly the formula above.
Only one term is ever active per example. For a positive example the loss is just −logp^:
| True label | Predicted p^ | Verdict | Loss −logp^ | Gradient at the logit (p^−y) |
|---|---|---|---|---|
| 1 | 0.90 | confident and right | 0.105 | −0.10 |
| 1 | 0.50 | no opinion | 0.693 | −0.50 |
| 1 | 0.10 | confidently wrong | 2.303 | −0.90 |
| 1 | 0.01 | catastrophically wrong | 4.605 | −0.99 |
The penalty is wildly asymmetric. Going from 0.9 to 0.5 costs 0.59; going from 0.1 to 0.01 costs another 2.3, and the loss heads to infinity as a confident prediction turns out flatly wrong. That punishes overconfidence far harder than hedging, which is what you want from a model whose probabilities you intend to trust. Compare the last column with the squared-error gradient above: at p^=0.01 cross-entropy delivers −0.99 against squared error's −0.0196. The sigmoid derivative that killed squared error cancels against a matching term in the cross-entropy derivative — fifty times more signal, exactly when it matters most.
The log-of-zero problem
If the model ever outputs exactly 0 or 1 — and in float32 a logit of just +17 already rounds to a sigmoid output of exactly 1 — you compute log(0)=−∞, the loss becomes inf, and every weight turns to nan on the next update. The run is dead and the traceback tells you nothing useful.
This is why frameworks give you a fused operation taking logits, not probabilities: BCEWithLogitsLoss in PyTorch, from_logits=True in Keras. Algebraically the whole expression simplifies to
which never takes a logarithm near zero and never exponentiates a large positive number. Same maths, no overflow. Applying a sigmoid yourself and then calling plain cross-entropy is one of the commonest causes of a run that silently produces nan.
If you are writing a sigmoid immediately before a cross-entropy, delete it and pass logits to the fused version instead.
More than two classes: softmax
With K classes the model emits K logits and softmax turns them into a distribution:
Exponentiating makes everything positive; dividing by the total makes them sum to 1. Categorical cross-entropy is L=−∑kyklogp^k, and since y is one-hot only the true class contributes: the loss is the negative log probability of the correct answer. Work it: logits (2.0,1.0,0.1) give ez=(7.389,2.718,1.105), summing to 11.212, so p^=(0.659,0.242,0.099). If class 0 is correct, the loss is −ln0.659=0.417.
Now the result worth memorising. Differentiate softmax-plus-cross-entropy with respect to the logits and almost everything cancels:
The gradient is simply predicted probability minus target. For our example, (−0.341,0.242,0.099): push the correct logit up, push the others down, in exact proportion to how much probability mass sits in the wrong place. No leftover softmax-derivative factor to shrink towards zero, almost free to compute, numerically stable. This cancellation is why softmax and cross-entropy are always fused into one operation.
Focal loss for lopsided data
Suppose 99% of your examples are negatives the model already gets right. Each contributes a small loss, but there are so many that together they drown out the rare positives you care about. Focal loss multiplies cross-entropy by a factor that collapses as confidence grows:
where p^t is the probability given to the true class and γ is typically 2.
| Example | p^t | Cross-entropy | Factor (1−p^t)2 | Focal loss |
|---|---|---|---|---|
| easy, already right | 0.90 | 0.105 | 0.01 | 0.00105 |
| hard, still wrong | 0.10 | 2.303 | 0.81 | 1.865 |
Under plain cross-entropy the hard example matters about 22 times more than the easy one; under focal loss, about 1,780 times more. Ten thousand easy negatives now add up to about as much loss as six hard positives, where under plain cross-entropy they outweighed more than four hundred. Note what focal loss keys on: not the class label, but the model's current confidence, recomputed every step.
Losses that score comparisons, not values
Sometimes there is no target value at all. In face verification or recommendation you want an embedding space where similar things sit close together, and "close" only means anything relative to other pairs.
Triplet loss takes an anchor a, a positive p that should be near it, and a negative n that should be far: ℓ=max(0,d(a,p)−d(a,n)+α). The margin α is the gap you insist on. With α=0.2: if d(a,p)=0.4 and d(a,n)=0.9, the bracket is −0.3, clipped to 0, so this triplet gives no gradient — it is comfortably correct, leave it alone. If d(a,p)=0.70 and d(a,n)=0.75, the bracket is 0.15: ordered correctly but only just, so the loss still pushes. Without the margin, a model could satisfy every triplet by a millionth of a unit and learn a fragile, collapsed embedding.
Contrastive loss works on pairs: ℓ=yd2+(1−y)max(0,m−d)2, with y=1 for "same". Similar pairs are pulled together with no lower limit; dissimilar pairs are pushed apart only until they clear the margin m, then ignored — otherwise the model spends its capacity shoving already-distant pairs further into the void.
Bolting a penalty onto the loss
A loss need not be purely about accuracy. A common pattern is
where R(θ) penalises the parameters themselves — often ∥θ∥22, the sum of squared weights. The optimiser now serves two masters and λ sets the exchange rate: at λ=0 the penalty vanishes; make it enormous and every weight is crushed to zero. The result is still one number, which is all the optimiser ever needed.
Computing them side by side
1import numpy as np23# --- Regression: one prediction has blown up ---4y = np.array([3.0, -0.5, 2.0, 7.0, 4.2])5y_hat = np.array([2.5, 0.0, 2.0, 8.0, 30.0])67def mse(y, p):8 return np.mean((y - p) ** 2)910def mae(y, p):11 return np.mean(np.abs(y - p))1213def huber(y, p, delta=1.0):14 a = np.abs(y - p)15 quad = np.minimum(a, delta) # inside the delta band16 lin = a - quad # outside it17 return np.mean(0.5 * quad ** 2 + delta * lin)1819print(mse(y, y_hat)) # 133.43 -> one row is 99.8% of this20print(mae(y, y_hat)) # 5.5621print(huber(y, y_hat)) # 5.21 -> barely moved by the outlier2223# --- Classification ---24y_cls = np.array([1, 1, 0, 0])25p = np.array([0.9, 0.6, 0.1, 0.4])2627def bce(y, p, eps=1e-12):28 p = np.clip(p, eps, 1 - eps) # the crude fix for log(0)29 return -np.mean(y * np.log(p) + (1 - y) * np.log(1 - p))3031def bce_from_logits(y, z):32 # max(z,0) - z*y + log(1 + exp(-|z|)): the fused, stable form33 return np.mean(np.maximum(z, 0) - z * y + np.log1p(np.exp(-np.abs(z))))3435def focal(y, p, gamma=2.0, eps=1e-12):36 p = np.clip(p, eps, 1 - eps)37 pt = np.where(y == 1, p, 1 - p) # probability of the TRUE class38 return -np.mean((1 - pt) ** gamma * np.log(pt))3940print(bce(y_cls, p)) # 0.308141print(focal(y_cls, p)) # 0.0414 -> confident examples nearly silenced4243z = np.log(p / (1 - p)) # recover the logits44assert np.isclose(bce(y_cls, p), bce_from_logits(y_cls, z))Huber and MAE land within 7% of each other while MSE is twenty-five times larger, all from a single bad row. And the focal number is far smaller than the cross-entropy number — which says nothing about model quality, only that a different function was applied.
Where people go wrong
| The belief | What is actually true |
|---|---|
| "The loss is the metric." | Different objects, different jobs. Accuracy, F1 and AUC are step functions with zero gradient almost everywhere, so they cannot be optimised directly. You train on a differentiable surrogate — cross-entropy — and report the metric. When they disagree, the metric is right about quality and the loss is right about what the model is doing. |
| "Training loss went down, so the model got better." | It means the model fits the training rows better, which any flexible enough model achieves perfectly by memorising them. Only held-out loss says anything about a new customer or tomorrow. |
| "Mine scores 0.03, yours scores 1.2, so mine wins." | Loss values compare only across models sharing the same loss, data and target scale. Cross-entropy of 0.03 and Huber of 1.2 are in incomparable units, and rescaling a target from pounds to thousands divides MSE by a million without touching the model. |
| "Class weights and focal loss do the same job." | Class weights multiply every example of a class by a fixed constant, whether the model finds it trivial or impossible. Focal loss weights by current confidence, so an easy positive is suppressed just like an easy negative. Class weights when imbalance is the problem; focal loss when a flood of easy examples is. They are often combined. |
Choose the loss before you choose the model
The practical discipline is to write down the cost of each kind of mistake before writing any code — in the units your organisation cares about, not in the abstract. What does being 20% over on a delivery estimate cost? What does 20% under cost? If those numbers differ, a symmetric loss is already wrong, and no amount of architecture search will fix it.
Then work backwards. Errors of both signs equally bad, data clean: mean squared error. Genuine outliers you do not want the model chasing: Huber. A typical case rather than an average case: mean absolute error, understanding that your predictions are now medians. Probabilities you will threshold or feed into a downstream decision: cross-entropy, fused with its sigmoid or softmax. A rare event buried under easy negatives: class weights, focal loss, or both.
Then run your chosen loss against a stupid baseline — predict the mean, the majority class, or last week's value — and read both numbers. A loss value alone is meaningless; a loss value beside a baseline is the first honest thing your model has told you. And when the model surprises you, resist blaming it: read the loss again. Nine times in ten it is doing exactly what you asked, and the surprise lives in the gap between what you wrote and what you meant.