Machine Learning Essentials

Regression Evaluation Metrics


Two models predict delivery times for a courier company. You have to ship one.

Model AModel B
Mean absolute error4.2 min6.8 min
Root mean squared error18.9 min11.2 min

Model A wins on one metric and loses on the other, and the disagreement is not a rounding artefact — it is a factual statement about how each model is wrong. Model A is closer on the typical delivery but occasionally catastrophically off. Model B is a bit worse most of the time and never wildly wrong.

Which one you ship depends entirely on something no metric can tell you: what a 45-minute error costs compared to three 15-minute errors. If customers cancel after 30 minutes of unexplained delay, Model A's rare disasters are the whole problem and Model B wins. If the number feeds a warehouse staffing model where small errors accumulate, Model A wins.

So the first thing to understand about regression metrics is that they are not different levels of accuracy. They are different definitions of "wrong", and choosing between them is a business decision wearing a mathematical costume.

Two courier models, and each wins on one metric4.2 min18.9 min4.56.8 min11.2 min1.6MAERMSERMSE / MAEA — rare disastersB — steadyA is usually closer; an RMSE four and a half times its MAE is the sign of rare, large misses.
Choosing RMSE tells the model that one 45-minute miss is worse than three 15-minute ones.

The raw material: residuals

Everything is built from residuals, ei=yi−y^ie_i = y_i - \hat{y}_i. Here are eight predictions from a house price model, in thousands of pounds:

ActualPredictedResidual|Residual|Residual²
200195+5525
250262−1212144
310298+1212144
180188−8864
420405+1515225
275270+5525
330337−7749
900420+480480230,400
Sum544231,076

The last row is a mansion in a dataset of ordinary houses — a genuine data point, badly predicted.

MAE =544/8=68= 544 / 8 = 68. MSE =231,076/8=28,884.5= 231{,}076 / 8 = 28{,}884.5. RMSE =28,884.5=170= \sqrt{28{,}884.5} = 170.

Now delete that one row and recompute over the remaining seven: MAE = 9.1, RMSE = 9.8. One row out of eight moved MAE by a factor of 7 and RMSE by a factor of 17.

MAE tells you about the typical error. RMSE tells you about the worst errors. When they diverge, you have outliers, and the size of the divergence measures how bad they are.

MAE and RMSE, precisely

MAE=1n∑i=1n∣yi−y^i∣RMSE=1n∑i=1n(yi−y^i)2\text{MAE} = \frac{1}{n}\sum_{i=1}^{n}|y_i - \hat{y}_i| \qquad \text{RMSE} = \sqrt{\frac{1}{n}\sum_{i=1}^{n}(y_i - \hat{y}_i)^2}

Both are in the units of the target, which is their great practical virtue — an RMSE of 170 means "typically off by about £170,000", a sentence anyone can act on. MSE, being in squared units, is fine as an optimisation objective and useless as a reported number: nobody knows what 28,884 squared thousand-pounds means.

A guaranteed relationship worth remembering: RMSE≥MAE\text{RMSE} \geq \text{MAE}, always. They are equal only if every error has exactly the same magnitude. The ratio RMSE/MAE is therefore a free outlier diagnostic:

RMSE / MAEInterpretation
≈ 1.0All errors similar size — unusual, often means a constant bias
1.1 – 1.5Healthy, roughly normal error distribution
1.5 – 3Meaningful heavy tail — some predictions much worse than typical
> 3A handful of rows dominate the error. Go and look at them individually

In the eight-row example the ratio is 170/68 = 2.5, and the diagnosis — go and look at the individual rows — leads straight to the mansion.

The deeper difference: what each one makes the model predict

This is the part most people never learn, and it explains far more than the outlier story.

Suppose for a given set of features, the true outcomes vary: identical-looking deliveries take 20, 22, 24, 25, and 60 minutes. What single number should the model output?

  • Training to minimise squared error produces the mean: (20+22+24+25+60)/5 = 30.2 minutes.
  • Training to minimise absolute error produces the median: 24 minutes.

These are not slightly different — 30.2 is longer than four of the five actual deliveries. Optimising MSE gives you a model that is systematically pulled towards rare large values, because that is mathematically the right answer when squared error is your definition of wrong.

So the choice of loss silently determines what quantity your model estimates. If you want a prediction that is "right most of the time", train on absolute error. If you want a prediction whose long-run total is right — because you will sum thousands of them for capacity planning — train on squared error.

R²: comparing against doing nothing

MAE and RMSE are in target units, which is helpful for interpretation and unhelpful for judgement. Is an RMSE of £28,000 good? Impossible to say without knowing whether houses cost £50,000 or £5 million.

R² fixes that by comparing your model to the dumbest possible one — always predicting the mean:

R2=1−∑i(yi−y^i)2∑i(yi−yˉ)2=1−your model’s squared errorthe mean-predictor’s squared errorR^2 = 1 - \frac{\sum_i (y_i - \hat{y}_i)^2}{\sum_i (y_i - \bar{y})^2} = 1 - \frac{\text{your model's squared error}}{\text{the mean-predictor's squared error}}

R²Meaning
1.0Perfect predictions. On real held-out data, suspect leakage
0.85Model explains 85% of the variance in the target
0.0Exactly as good as always guessing the mean
−0.4Worse than guessing the mean. Possible and common on test sets

Negative R² on a test set surprises people who half-remember that "R² is between 0 and 1". That range applies to the training data of an ordinary least-squares fit. On new data a bad model can absolutely do worse than a constant, and R² correctly reports it.

Three traps in R²

It always increases when you add features. Add a column of random noise and training R² will tick up, because the model can use it to fit a little more of the training noise. This makes R² useless for choosing between models with different numbers of features unless you use the adjusted version:

Radj2=1−(1−R2)n−1n−p−1R^2_{\text{adj}} = 1 - (1 - R^2)\frac{n-1}{n-p-1}

With n=100n = 100 rows and p=30p = 30 features, an R² of 0.80 becomes an adjusted R² of 0.71. The penalty grows as features approach sample size, which is exactly when unadjusted R² becomes most misleading.

It depends on the spread of your test set, not just your model. The denominator is the variance of the actual values. Evaluate the same model on a homogeneous neighbourhood where all houses cost £280,000–£320,000 and R² will collapse, because there was little variance to explain. Evaluate it on a mixed city and R² looks excellent. The model did not change. This makes R² dangerous for comparing performance across different datasets or segments.

High R² does not mean the model is correct. A model with R² of 0.92 whose residuals form a clear curve is a wrong model that happens to capture the broad trend. The metric cannot see structure in the errors; only a plot can.

Percentage errors, and why they misbehave

MAPE=100n∑i=1n∣yi−y^iyi∣\text{MAPE} = \frac{100}{n}\sum_{i=1}^{n}\left|\frac{y_i - \hat{y}_i}{y_i}\right|

MAPE is scale-free and communicates beautifully: "our forecasts are off by 8% on average" needs no further explanation. It has three failure modes, and all three are common.

It is undefined when the actual value is zero, and explodes when it is near zero. Forecast 3 units of demand when the truth is 1 and MAPE records 200%. Forecast 300 when the truth is 305 and it records 1.6%. In demand forecasting with many low-volume items, MAPE is dominated entirely by the small ones.

It is asymmetric. Over-forecasting is penalised more heavily than under-forecasting. Actual 100, predicted 150 gives 50%. Actual 100, predicted 50 also gives 50% — but the ceiling differs: under-forecasting can never exceed 100% error, while over-forecasting is unbounded. A model tuned on MAPE learns to under-forecast systematically.

It implies a cost model you may not hold. MAPE says a £10 error on a £100 item is as bad as a £1,000 error on a £10,000 item. Sometimes true, often not.

The usual alternatives:

MetricFormula sketchUse when
MAPEmean of ∣e∣/y|e|/yAll actuals comfortably above zero and similar in scale
sMAPE∣e∣|e| over the average of actual and predictedYou need symmetry; bounded at 200%
MedAEmedian of ∣e∣|e|Outliers exist and should be ignored entirely
RMSLERMSE computed on log⁡(1+y)\log(1+y)Target spans orders of magnitude; relative error matters; under-prediction is worse
MASEMAE divided by a naive forecast's MAETime series; you want a scale-free number that is honest near zero

RMSLE deserves a note because it appears constantly in problems with skewed targets like house prices, sales, and view counts. By taking logs first, it measures ratio error rather than absolute error: predicting 200 when the truth is 100 costs the same as predicting 2,000 when the truth is 1,000. It also penalises under-prediction more than over-prediction, which suits inventory problems where running out is worse than overstocking.

Choosing, without agonising

Your situationPrimary metricWhy
Large errors are disproportionately costlyRMSESquaring makes the model fear them
All errors cost proportionally to sizeMAELinear cost, linear metric
Data contains genuine outliers you cannot fixMAE or MedAEStops a few rows dictating the model
Target spans orders of magnitudeRMSLE or MAPERelative error is the meaningful quantity
Explaining to a non-technical stakeholderMAE plus R²"Off by 12 minutes; explains 78% of the variation"
Comparing across different datasetsR² or MASEScale-free
Predictions get summed downstreamRMSE, and check bias separatelySystematic bias compounds when summed

Report at least two. A single metric hides too much, and the disagreement between two metrics is itself information — as the opening comparison showed.

One metric nobody reports and everybody should

Mean error, without the absolute value:

Bias=1n∑i(yi−y^i)\text{Bias} = \frac{1}{n}\sum_i (y_i - \hat{y}_i)

This is near zero for a well-calibrated model. If it comes out at +£14,000, your model systematically underpredicts by £14,000, and that is a fixable problem invisible to MAE, RMSE, and R², all of which discard the sign. Systematic bias matters enormously whenever predictions are aggregated: 10,000 deliveries each underestimated by 3 minutes is 500 hours of unplanned labour.

Computing them properly

Python
import numpy as npfrom sklearn.metrics import (mean_absolute_error, mean_squared_error,                             r2_score, median_absolute_error,                             mean_absolute_percentage_error)def report(y_true, y_pred, label=""):    resid = y_true - y_pred    mae = mean_absolute_error(y_true, y_pred)    rmse = np.sqrt(mean_squared_error(y_true, y_pred))    print(f"--- {label} ---")    print(f"MAE      {mae:12,.2f}")    print(f"RMSE     {rmse:12,.2f}")    print(f"RMSE/MAE {rmse / mae:12.2f}   (outlier indicator)")    print(f"MedAE    {median_absolute_error(y_true, y_pred):12,.2f}")    print(f"MAPE     {mean_absolute_percentage_error(y_true, y_pred) * 100:11,.2f}%")    print(f"R2       {r2_score(y_true, y_pred):12.4f}")    print(f"Bias     {resid.mean():12,.2f}   (should be near 0)")    print(f"P90 |err|{np.percentile(np.abs(resid), 90):12,.2f}")report(y_test, model.predict(X_test), "gradient boosting")

The 90th percentile of absolute error is worth including. It answers "how bad is a bad day?" directly, without the interpretive step RMSE requires.

The sign convention that confuses everyone

Scikit-learn's cross_val_score follows a convention that higher is better, so error metrics are returned negated:

Python
from sklearn.model_selection import cross_val_scorescores = cross_val_score(model, X, y, cv=5,                         scoring="neg_root_mean_squared_error")print(f"RMSE: {-scores.mean():,.0f} +/- {scores.std():,.0f}")

Forget the minus sign and you will report an RMSE of −28,412, which is not a small error but a sign error.

Metrics tell you how much; residuals tell you why

No aggregate number can tell you where a model fails. For that, plot the residuals against the predictions.

Python
import matplotlib.pyplot as pltpred = model.predict(X_test)resid = y_test - predfig, ax = plt.subplots(1, 3, figsize=(15, 4))ax[0].scatter(pred, resid, alpha=0.4); ax[0].axhline(0, c="k")ax[0].set_xlabel("predicted"); ax[0].set_ylabel("residual")ax[1].hist(resid, bins=40)ax[2].scatter(y_test, pred, alpha=0.4)lims = [y_test.min(), y_test.max()]ax[2].plot(lims, lims, "k--")     # perfect prediction line
Pattern in the residual plotDiagnosisAction
Shapeless cloud around zeroModel has captured the structureNothing — this is the goal
Curve or U shapeMissing nonlinearityPolynomial terms, or a tree-based model
Funnel widening to the rightError grows with the targetModel log⁡(y)\log(y); consider RMSLE
Residuals all positive at high predictionsModel compresses the rangeUnder-fitting at the extremes; add capacity
Two separate bandsA missing categorical featureFind and add the group variable
A few points far from the restOutliers or data errorsInspect those rows individually

The compression pattern is worth recognising because it is nearly universal in tree-based models. Because a tree predicts the average of a leaf, it can never predict beyond the range it saw, so expensive houses are underpredicted and cheap ones overpredicted. In the actual-versus-predicted plot this shows as points falling on a line flatter than the diagonal.

Segment before you conclude

An RMSE of £28,000 across all houses can hide an RMSE of £12,000 on ordinary homes and £145,000 on the top 3%. If your business only cares about ordinary homes, the aggregate number has understated your model by a factor of two. If it cares about the expensive ones, the aggregate has hidden a total failure.

Python
import pandas as pdresults = pd.DataFrame({"actual": y_test, "pred": pred})results["abs_err"] = (results.actual - results.pred).abs()results["band"] = pd.qcut(results.actual, 5,                          labels=["lowest", "low", "mid", "high", "highest"])print(results.groupby("band", observed=True)      .agg(n=("abs_err", "size"),           mae=("abs_err", "mean"),           bias=("actual", lambda s: (s - results.loc[s.index, "pred"]).mean())))

Do this by price band, by region, by month, and by any segment a stakeholder cares about. Aggregate metrics are averages, and averages conceal precisely the failures that get a model pulled from production.

What this means when you build something

Decide the metric before you train anything, in a conversation with whoever will use the output, and phrase the question in their language: "if we are 30 minutes late once a week, is that worse or better than being 8 minutes late every day?" Their answer chooses between RMSE and MAE more reliably than any rule of thumb.

Then report three things rather than one: a scale metric in the units people think in (MAE), a robustness check (RMSE/MAE ratio, or the 90th percentile error), and a comparison against doing nothing (R², or the error of a mean-predictor baseline). Those three together are almost impossible to misread.

And never let the numbers be the last thing you look at. Plot the residuals. A model with excellent metrics and a curved residual plot is telling you there is a pattern still sitting in your errors — free accuracy you have not collected yet.