Mathematics for Machine Learning

Loss Functions — Teaching a Model What "Wrong" Means


Two models predict house prices. Model A is off by £5,000 on every single house. Model B is exactly right on 90 houses out of 100 and off by £200,000 on the other ten. Which model is better?

You cannot answer that. Not because you lack information — you have every prediction and every true price — but because "better" is not a property of the numbers. It is a decision you have to make and write down. A mortgage lender pricing risk and an estate agent writing brochures would answer differently, and both would be right.

Try the obvious rule: count how many predictions are exactly correct. Model B wins, 90 to nothing. But that rule is useless for training. With continuous outputs you almost never hit a value exactly, so the count is zero for nearly every model you could build, and being off by £1 scores the same as being off by £1 million. Worse, the count is a staircase: nudge a weight by a hair and it does not move. Gradient-based learning needs a slope to slide down, and a staircase has none.

So you need something else: a function that eats one true value and one prediction and returns a single number saying how bad that was, smoothly enough to have a slope. That is a loss function. And here is what catches people out: the model does not care what you meant. It will push that number down with total, literal obedience.

What squaring the residual asks the model to doSquared error• A miss of 10 costs 100 times a miss of 1• Fits the conditional mean• One outlier can steer the whole fit• Smooth everywhere, easy to differentiateAbsolute error• A miss of 10 costs 10 times a miss of 1• Fits the conditional median• Outliers get no extra vote• Gradient is undefined at exactly zero
The loss is not a scoreboard — it is the instruction telling the model which quantity to predict.

Two levels: one example, and the whole dataset

The per-example loss ℓ(y,y^)\ell(y, \hat y) scores a single prediction: non-negative, zero when perfect, larger as the prediction gets worse. It knows nothing about the model or the rest of the data — just one truth and one guess.

The empirical risk is that averaged over your training set:

L(θ)=1n∑i=1nℓ(yi,y^i(θ))L(\theta) = \frac{1}{n}\sum_{i=1}^{n} \ell\big(y_i, \hat y_i(\theta)\big)

Here θ\theta stands for every parameter you are allowed to change, and training means finding the θ\theta that makes L(θ)L(\theta) smallest. The word empirical matters: this is the loss on the finite pile of data you happen to have, standing in for the loss you actually care about, the average over all the data you will never see. That gap is why a model can look brilliant in training and fall over in production.

Dividing by nn is not cosmetic. Without it, doubling your dataset doubles every gradient, so a learning rate that worked on 10,000 rows explodes on 20,000.

A model does not learn the task you had in mind. It learns to make one number small. Everything else is your responsibility.

What makes a loss usable

  • Differentiable almost everywhere. Isolated kinks are fine — like ∣x∣|x| at zero — flat regions are not. Accuracy fails badly: it is a step function whose derivative is zero wherever it exists, so the optimiser is told "moving does nothing" and stands still.
  • Aligned with the goal. The loss is the only channel through which your intentions reach the model. If a missed fraud costs £10,000 and a false alarm costs a phone call, and the loss treats both identically, the model will happily trade fraud for phone calls.
  • Well-scaled. Loss values around 10610^6 produce enormous gradients and demand an absurdly small learning rate; values around 10−810^{-8} vanish into floating-point noise. Predicting prices in pounds rather than thousands multiplies squared-error loss values by a million and their gradients by a thousand.
  • Convex in the prediction, ideally. Convex means bowl-shaped: one bottom, no local traps. A deep network destroys global convexity anyway, but a convex per-prediction loss guarantees the loss is not adding minima of its own on top of the network's.

Regression losses: the choice is really about outliers

Mean squared error

MSE=1n∑i=1n(yi−y^i)2,∂L∂y^i=2n(y^i−yi)\text{MSE} = \frac{1}{n}\sum_{i=1}^{n}(y_i - \hat y_i)^2, \qquad \frac{\partial L}{\partial \hat y_i} = \frac{2}{n}(\hat y_i - y_i)

Square each error so overshooting by 3 and undershooting by 3 cost the same, then average. Read the gradient in words: the push each example applies is proportional to how wrong it is. An example off by 10 pulls ten times harder than one off by 1. On clean data that is reasonable. On dirty data it is a trap.

The concrete cost of squaring

Take ten predictions. Nine are off by exactly 1. One — a mistyped data entry, say — is off by 10.

MSE: the nine contribute 9×12=99 \times 1^2 = 9, the bad one contributes 102=10010^2 = 100. Total 109, so MSE=10.9\text{MSE} = 10.9. The single outlier is 91.7% of the entire loss.

MAE (mean absolute error, 1n∑∣yi−y^i∣\frac{1}{n}\sum|y_i - \hat y_i|): the nine contribute 9, the outlier contributes 10, so MAE=1.9\text{MAE} = 1.9. The outlier is 52.6% of the loss.

Same data, same predictions. Under MSE nine-tenths of the loss comes from one possibly-corrupt row, and because the MSE gradient grows with the error, that row also supplies more than half of the push on the model (10 of the 19 units of error). The model bends itself out of shape to accommodate it. Under MAE the row counts as ten ordinary rows in the loss, but its push is the same ±1\pm 1 as every other row's — heavy, but not dictatorial.

Mean absolute error and its awkward corner

MAE's gradient is ±1\pm 1 regardless of the size of the error: only the direction changes. That is what makes it robust and also what makes it awkward. As you converge the gradient never shrinks, so the optimiser keeps taking full-size steps and bounces around the minimum instead of settling into it; you need a decaying learning rate to compensate.

MAE is also not differentiable at zero, the point where the prediction is exactly right. Formally you use a subgradient — any slope between −1-1 and +1+1 is valid there. Practically every framework returns 0 and nothing bad happens, because landing on an exact tie is vanishingly rare in floating point.

Huber loss: the practical compromise

Writing the error as a=y−y^a = y - \hat y:

ℓδ(a)={12a2if ∣a∣≤δδ(∣a∣−12δ)otherwise\ell_\delta(a) = \begin{cases} \tfrac{1}{2}a^2 & \text{if } |a| \le \delta \\[4pt] \delta\left(|a| - \tfrac{1}{2}\delta\right) & \text{otherwise}\end{cases}

Quadratic for small errors, linear for large ones. The −12δ-\tfrac{1}{2}\delta is not decoration: it makes the two pieces meet at ∣a∣=δ|a| = \delta with the same value and the same slope, so there is no jump and no kink. The gradient rises with the error up to a magnitude of δ\delta, then stays capped.

So Huber behaves like MSE near the answer, where you want smooth convergence, and like MAE out in the tails, where you want protection. The parameter δ\delta is the line you draw between "ordinary error" and "outlier"; a common starting point is the 90th percentile of absolute residuals from a quick first fit. Large δ\delta recovers MSE, small δ\delta a scaled MAE.

LossPer-example formGradient behaviourOutlier sensitivityDifferentiablePulls prediction towards
MSE(y−y^)2(y-\hat y)^2Grows with the error, unboundedVery high — errors are squaredEverywherethe mean
MAE∣y−y^∣|y-\hat y|Constant ±1\pm 1, never shrinksLow — errors counted onceEverywhere except y^=y\hat y = ythe median
Huber(δ\delta)quadratic inside δ\delta, linear outsideGrows to ±δ\pm\delta, then cappedLow, and smooth near zeroEverywhere, including the joinsbetween the two

Why the last column is the real reason to choose

Ask what single constant cc minimises each loss on a fixed set of numbers. Differentiate ∑i(yi−c)2\sum_i (y_i - c)^2 and set to zero: −2∑i(yi−c)=0-2\sum_i(y_i - c) = 0, giving c=1n∑yic = \frac{1}{n}\sum y_i, the mean. The derivative of ∑i∣yi−c∣\sum_i |y_i - c| counts how many points sit above cc minus how many sit below; that is zero when the counts match, which is the median.

On the data 2,4,5,6,1002, 4, 5, 6, 100: MSE drives your prediction towards 23.423.4, MAE towards 55. Neither is wrong. Forecasting warehouse stock you will later sum? You need unbiased totals, so you need the mean, so you need MSE. Telling a customer how long delivery takes and wanting to be right half the time? That is the median, so MAE.

Choosing between squared and absolute error is not a technicality about outliers. It decides whether your model reports the mean or the median.

Classification: why squared error is the wrong tool

The obvious move is to encode the label as 0 or 1, output a probability p^\hat p, and use (y−p^)2(y - \hat p)^2. It fails for two reasons.

A classifier produces a raw score zz (a logit) and squashes it with the sigmoid p^=1/(1+e−z)\hat p = 1/(1 + e^{-z}), whose derivative is p^(1−p^)\hat p(1 - \hat p). By the chain rule, ∂∂z(p^−y)2=2(p^−y) p^(1−p^)\frac{\partial}{\partial z}(\hat p - y)^2 = 2(\hat p - y)\,\hat p(1-\hat p). Suppose the true label is 1 and the model confidently says p^=0.01\hat p = 0.01 — as wrong as it gets. Then p^(1−p^)=0.0099\hat p(1-\hat p) = 0.0099 and the gradient is 2×(−0.99)×0.0099≈−0.01962 \times (-0.99) \times 0.0099 \approx -0.0196. Tiny. The model is maximally wrong and gets almost no correction, because the sigmoid has saturated: learning stalls exactly where it is most needed. Second, squared error composed with a sigmoid is not convex even for a single-layer model, so you inherit local minima you did not need.

Binary cross-entropy, and where it comes from

L=−1n∑i=1n[yilog⁡p^i+(1−yi)log⁡(1−p^i)]L = -\frac{1}{n}\sum_{i=1}^{n}\Big[y_i \log \hat p_i + (1-y_i)\log(1 - \hat p_i)\Big]

This is not arbitrary; it falls out of maximum likelihood. Treat each label as a coin flip whose bias the model predicts. The probability the model assigns to the label actually observed is p^ y(1−p^)1−y\hat p^{\,y}(1-\hat p)^{1-y} — check it: y=1y=1 gives p^\hat p, y=0y=0 gives 1−p^1-\hat p. Assume independence, so the dataset's probability is the product of those terms. Products of thousands of numbers below 1 underflow to zero, so take the logarithm, turning the product into a sum; maximising is minimising the negative, so flip the sign; divide by nn. That is exactly the formula above.

Only one term is ever active per example. For a positive example the loss is just −log⁡p^-\log \hat p:

True labelPredicted p^\hat pVerdictLoss −log⁡p^-\log \hat pGradient at the logit (p^−y\hat p - y)
10.90confident and right0.105−0.10-0.10
10.50no opinion0.693−0.50-0.50
10.10confidently wrong2.303−0.90-0.90
10.01catastrophically wrong4.605−0.99-0.99

The penalty is wildly asymmetric. Going from 0.9 to 0.5 costs 0.59; going from 0.1 to 0.01 costs another 2.3, and the loss heads to infinity as a confident prediction turns out flatly wrong. That punishes overconfidence far harder than hedging, which is what you want from a model whose probabilities you intend to trust. Compare the last column with the squared-error gradient above: at p^=0.01\hat p = 0.01 cross-entropy delivers −0.99-0.99 against squared error's −0.0196-0.0196. The sigmoid derivative that killed squared error cancels against a matching term in the cross-entropy derivative — fifty times more signal, exactly when it matters most.

The log-of-zero problem

If the model ever outputs exactly 0 or 1 — and in float32 a logit of just +17+17 already rounds to a sigmoid output of exactly 1 — you compute log⁡(0)=−∞\log(0) = -\infty, the loss becomes inf, and every weight turns to nan on the next update. The run is dead and the traceback tells you nothing useful.

This is why frameworks give you a fused operation taking logits, not probabilities: BCEWithLogitsLoss in PyTorch, from_logits=True in Keras. Algebraically the whole expression simplifies to

max⁡(z,0)−z y+log⁡ ⁣(1+e−∣z∣)\max(z, 0) - z\,y + \log\!\left(1 + e^{-|z|}\right)

which never takes a logarithm near zero and never exponentiates a large positive number. Same maths, no overflow. Applying a sigmoid yourself and then calling plain cross-entropy is one of the commonest causes of a run that silently produces nan.

If you are writing a sigmoid immediately before a cross-entropy, delete it and pass logits to the fused version instead.

More than two classes: softmax

With KK classes the model emits KK logits and softmax turns them into a distribution:

p^k=ezk∑j=1Kezj\hat p_k = \frac{e^{z_k}}{\sum_{j=1}^{K} e^{z_j}}

Exponentiating makes everything positive; dividing by the total makes them sum to 1. Categorical cross-entropy is L=−∑kyklog⁡p^kL = -\sum_k y_k \log \hat p_k, and since yy is one-hot only the true class contributes: the loss is the negative log probability of the correct answer. Work it: logits (2.0, 1.0, 0.1)(2.0,\, 1.0,\, 0.1) give ez=(7.389, 2.718, 1.105)e^z = (7.389,\, 2.718,\, 1.105), summing to 11.21211.212, so p^=(0.659, 0.242, 0.099)\hat p = (0.659,\, 0.242,\, 0.099). If class 0 is correct, the loss is −ln⁡0.659=0.417-\ln 0.659 = 0.417.

Now the result worth memorising. Differentiate softmax-plus-cross-entropy with respect to the logits and almost everything cancels:

∂L∂zk=p^k−yk\frac{\partial L}{\partial z_k} = \hat p_k - y_k

The gradient is simply predicted probability minus target. For our example, (−0.341, 0.242, 0.099)(-0.341,\, 0.242,\, 0.099): push the correct logit up, push the others down, in exact proportion to how much probability mass sits in the wrong place. No leftover softmax-derivative factor to shrink towards zero, almost free to compute, numerically stable. This cancellation is why softmax and cross-entropy are always fused into one operation.

Focal loss for lopsided data

Suppose 99% of your examples are negatives the model already gets right. Each contributes a small loss, but there are so many that together they drown out the rare positives you care about. Focal loss multiplies cross-entropy by a factor that collapses as confidence grows:

ℓ=−(1−p^t)γlog⁡p^t\ell = -(1 - \hat p_t)^{\gamma}\log \hat p_t

where p^t\hat p_t is the probability given to the true class and γ\gamma is typically 2.

Examplep^t\hat p_tCross-entropyFactor (1−p^t)2(1-\hat p_t)^2Focal loss
easy, already right0.900.1050.010.00105
hard, still wrong0.102.3030.811.865

Under plain cross-entropy the hard example matters about 22 times more than the easy one; under focal loss, about 1,780 times more. Ten thousand easy negatives now add up to about as much loss as six hard positives, where under plain cross-entropy they outweighed more than four hundred. Note what focal loss keys on: not the class label, but the model's current confidence, recomputed every step.

Losses that score comparisons, not values

Sometimes there is no target value at all. In face verification or recommendation you want an embedding space where similar things sit close together, and "close" only means anything relative to other pairs.

Triplet loss takes an anchor aa, a positive pp that should be near it, and a negative nn that should be far: ℓ=max⁡(0,  d(a,p)−d(a,n)+α)\ell = \max\big(0,\; d(a,p) - d(a,n) + \alpha\big). The margin α\alpha is the gap you insist on. With α=0.2\alpha = 0.2: if d(a,p)=0.4d(a,p) = 0.4 and d(a,n)=0.9d(a,n) = 0.9, the bracket is −0.3-0.3, clipped to 0, so this triplet gives no gradient — it is comfortably correct, leave it alone. If d(a,p)=0.70d(a,p) = 0.70 and d(a,n)=0.75d(a,n) = 0.75, the bracket is 0.150.15: ordered correctly but only just, so the loss still pushes. Without the margin, a model could satisfy every triplet by a millionth of a unit and learn a fragile, collapsed embedding.

Contrastive loss works on pairs: ℓ=y d2+(1−y)max⁡(0,  m−d)2\ell = y\,d^2 + (1-y)\max(0,\; m - d)^2, with y=1y=1 for "same". Similar pairs are pulled together with no lower limit; dissimilar pairs are pushed apart only until they clear the margin mm, then ignored — otherwise the model spends its capacity shoving already-distant pairs further into the void.

Bolting a penalty onto the loss

A loss need not be purely about accuracy. A common pattern is

Ltotal(θ)=1n∑iℓ(yi,y^i)⏟fit the data+λ R(θ)⏟stay simpleL_{\text{total}}(\theta) = \underbrace{\frac{1}{n}\sum_i \ell(y_i, \hat y_i)}_{\text{fit the data}} + \underbrace{\lambda\, R(\theta)}_{\text{stay simple}}

where R(θ)R(\theta) penalises the parameters themselves — often ∥θ∥22\|\theta\|_2^2, the sum of squared weights. The optimiser now serves two masters and λ\lambda sets the exchange rate: at λ=0\lambda = 0 the penalty vanishes; make it enormous and every weight is crushed to zero. The result is still one number, which is all the optimiser ever needed.

Computing them side by side

Python
import numpy as np# --- Regression: one prediction has blown up ---y     = np.array([3.0, -0.5, 2.0,  7.0,  4.2])y_hat = np.array([2.5,  0.0, 2.0,  8.0, 30.0])def mse(y, p):    return np.mean((y - p) ** 2)def mae(y, p):    return np.mean(np.abs(y - p))def huber(y, p, delta=1.0):    a    = np.abs(y - p)    quad = np.minimum(a, delta)      # inside the delta band    lin  = a - quad                  # outside it    return np.mean(0.5 * quad ** 2 + delta * lin)print(mse(y, y_hat))    # 133.43  -> one row is 99.8% of thisprint(mae(y, y_hat))    #   5.56print(huber(y, y_hat))  #   5.21  -> barely moved by the outlier# --- Classification ---y_cls = np.array([1, 1, 0, 0])p     = np.array([0.9, 0.6, 0.1, 0.4])def bce(y, p, eps=1e-12):    p = np.clip(p, eps, 1 - eps)     # the crude fix for log(0)    return -np.mean(y * np.log(p) + (1 - y) * np.log(1 - p))def bce_from_logits(y, z):    # max(z,0) - z*y + log(1 + exp(-|z|)): the fused, stable form    return np.mean(np.maximum(z, 0) - z * y + np.log1p(np.exp(-np.abs(z))))def focal(y, p, gamma=2.0, eps=1e-12):    p  = np.clip(p, eps, 1 - eps)    pt = np.where(y == 1, p, 1 - p)  # probability of the TRUE class    return -np.mean((1 - pt) ** gamma * np.log(pt))print(bce(y_cls, p))     # 0.3081print(focal(y_cls, p))   # 0.0414  -> confident examples nearly silencedz = np.log(p / (1 - p))                        # recover the logitsassert np.isclose(bce(y_cls, p), bce_from_logits(y_cls, z))

Huber and MAE land within 7% of each other while MSE is twenty-five times larger, all from a single bad row. And the focal number is far smaller than the cross-entropy number — which says nothing about model quality, only that a different function was applied.

Where people go wrong

The beliefWhat is actually true
"The loss is the metric."Different objects, different jobs. Accuracy, F1 and AUC are step functions with zero gradient almost everywhere, so they cannot be optimised directly. You train on a differentiable surrogate — cross-entropy — and report the metric. When they disagree, the metric is right about quality and the loss is right about what the model is doing.
"Training loss went down, so the model got better."It means the model fits the training rows better, which any flexible enough model achieves perfectly by memorising them. Only held-out loss says anything about a new customer or tomorrow.
"Mine scores 0.03, yours scores 1.2, so mine wins."Loss values compare only across models sharing the same loss, data and target scale. Cross-entropy of 0.03 and Huber of 1.2 are in incomparable units, and rescaling a target from pounds to thousands divides MSE by a million without touching the model.
"Class weights and focal loss do the same job."Class weights multiply every example of a class by a fixed constant, whether the model finds it trivial or impossible. Focal loss weights by current confidence, so an easy positive is suppressed just like an easy negative. Class weights when imbalance is the problem; focal loss when a flood of easy examples is. They are often combined.

Choose the loss before you choose the model

The practical discipline is to write down the cost of each kind of mistake before writing any code — in the units your organisation cares about, not in the abstract. What does being 20% over on a delivery estimate cost? What does 20% under cost? If those numbers differ, a symmetric loss is already wrong, and no amount of architecture search will fix it.

Then work backwards. Errors of both signs equally bad, data clean: mean squared error. Genuine outliers you do not want the model chasing: Huber. A typical case rather than an average case: mean absolute error, understanding that your predictions are now medians. Probabilities you will threshold or feed into a downstream decision: cross-entropy, fused with its sigmoid or softmax. A rare event buried under easy negatives: class weights, focal loss, or both.

Then run your chosen loss against a stupid baseline — predict the mean, the majority class, or last week's value — and read both numbers. A loss value alone is meaningless; a loss value beside a baseline is the first honest thing your model has told you. And when the model surprises you, resist blaming it: read the loss again. Nine times in ten it is doing exactly what you asked, and the surprise lives in the gap between what you wrote and what you meant.