Course Content
Machine Learning Essentials
6 sections · 16 lessons
Classification Evaluation Metrics
A screening model for a rare cancer reports 99.2% accuracy on 10,000 patients. The team is delighted. Then someone asks how many of the 80 actual cancer cases it found.
The answer is zero. The model predicts "no cancer" for every single patient. With 9,920 healthy people and 80 sick ones, that strategy is right 9,920 times out of 10,000 — 99.2% — and it is worth nothing, because a screening tool that never screens positive is a piece of paper that says "you're fine".
Accuracy is one number describing a system that fails in at least two distinct ways, and it weights those failures by how often they happen rather than by what they cost. To evaluate a classifier you need to separate the kinds of error, and everything useful starts from one small table.
The confusion matrix
Four numbers, and every classification metric is arithmetic on them.
| Predicted: cancer | Predicted: healthy | |
|---|---|---|
| Actually cancer | True Positive (TP) | False Negative (FN) — missed case |
| Actually healthy | False Positive (FP) — false alarm | True Negative (TN) |
The naming is easier than it looks: the second word is what the model said, the first says whether it was right. "False negative" = model said negative, and that was false.
A model that actually tries produces something like this:
| Predicted cancer | Predicted healthy | Row total | |
|---|---|---|---|
| Actually cancer | TP = 64 | FN = 16 | 80 |
| Actually healthy | FP = 310 | TN = 9,610 | 9,920 |
Accuracy is now (64+9610)/10000=96.7% — lower than the useless model's 99.2%. That single fact should retire accuracy from any imbalanced problem permanently.
Three metrics, three questions
Each takes a different slice of that table, and the trick to remembering them is to notice which total sits in the denominator.
Precision: of the alarms we raised, how many were real?
Denominator: the predicted positive column. 17% of flagged patients actually have cancer; 83% get a frightening letter and an unnecessary biopsy. Precision is the metric that matters when acting on a positive is expensive.
Recall: of the real cases, how many did we catch?
Denominator: the actually positive row. We found 80% of cancers and missed 16 people. Recall is the metric that matters when missing a positive is expensive.
Specificity: of the real negatives, how many did we correctly clear?
96.9% sounds superb, and it illustrates why specificity is a poor guide under imbalance. A 3.1% false-positive rate applied to 9,920 healthy people generates 310 false alarms — nearly four times the 80 real cases. Small percentages of a large group swamp large percentages of a small one.
Precision is judged by the column you predicted. Recall is judged by the row that was true. Confusing them is the single most common error in reporting classifier performance.
Why you cannot have both
Precision and recall trade off because both are controlled by the same knob: the probability threshold at which you declare a positive.
| Threshold | TP | FP | FN | Precision | Recall | F1 |
|---|---|---|---|---|---|---|
| 0.90 | 18 | 4 | 62 | 0.818 | 0.225 | 0.353 |
| 0.70 | 36 | 28 | 44 | 0.563 | 0.450 | 0.500 |
| 0.50 | 52 | 96 | 28 | 0.351 | 0.650 | 0.456 |
| 0.30 | 64 | 310 | 16 | 0.171 | 0.800 | 0.282 |
| 0.10 | 76 | 1,840 | 4 | 0.040 | 0.950 | 0.076 |
| 0.01 | 80 | 9,920 | 0 | 0.008 | 1.000 | 0.016 |
Raise the threshold and you only flag cases you are sure about: precision up, recall down. Lower it and you catch nearly everything at the cost of drowning in false alarms. The bottom row is the degenerate extreme — flag everyone, achieve perfect recall, and have said nothing.
The single most useful consequence: the same trained model produces every row of that table. Nobody needs to retrain to move along it. Yet teams routinely report the 0.5 row as though it were a property of the model rather than an arbitrary default.
F1 and its relatives
When you need one number, F1 is the harmonic mean of precision and recall:
The harmonic mean matters. With precision 0.9 and recall 0.1, the arithmetic mean is 0.50 — respectable-looking. F1 is 0.18, which correctly reflects that the model is nearly useless. The harmonic mean is dragged down by whichever value is smaller, so you cannot game it by maximising one side.
F1 assumes precision and recall are equally important, which they almost never are. Fβ lets you weight them:
| β | Meaning | Fits |
|---|---|---|
| 0.5 | Precision counts double | Spam filtering — a lost real email is worse than a spam that gets through |
| 1 | Equal weight | The default when you have no information |
| 2 | Recall counts double | Disease screening, fraud — a miss is worse than a false alarm |
ROC, AUC, and where AUC lies to you
The ROC curve plots true positive rate (recall) against false positive rate as the threshold sweeps from 1 to 0. AUC is the area beneath it.
AUC has an exact and useful interpretation, worth stating precisely: AUC is the probability that a randomly chosen positive case receives a higher score than a randomly chosen negative case. AUC 0.5 means the ranking is random. AUC 0.85 means that 85% of the time, a randomly picked sick patient scores above a randomly picked healthy one.
That framing shows both its strength and its weakness. It measures ranking quality and nothing else. It is threshold-independent, which is genuinely useful when you have not decided the threshold. It is also insensitive to imbalance in a way that flatters bad models.
The same model, two curves
Take the 0.30-threshold row above: 64 TP, 310 FP, 16 FN, 9,610 TN.
- True positive rate = 64/80 = 0.800
- False positive rate = 310/9,920 = 0.031
On a ROC curve that is a point near the top-left corner — an excellent-looking operating point, and the full curve here would give an AUC around 0.94.
On a precision-recall curve the same operating point is recall 0.800, precision 0.171: far down the chart. Both descriptions are correct. The difference is that the ROC's false positive rate divides by 9,920, so 310 false alarms look like a rounding error, while precision divides by 374, so they dominate.
| ROC / AUC | Precision-Recall / average precision | |
|---|---|---|
| Axes | Recall vs false positive rate | Precision vs recall |
| Uses true negatives? | Yes | No |
| Baseline for a random model | Always 0.5 | The positive class rate (here 0.008) |
| Under heavy imbalance | Optimistic | Honest |
| Best used when | Classes roughly balanced; you care about ranking | Positives are rare and are what you care about |
Rule of thumb: if the positive class is under about 10% of the data, report average precision alongside AUC, and treat a large gap between them as the imbalance showing through.
Multi-class: the averaging choice changes the story
With more than two classes, per-class precision and recall are computed one-vs-rest, then averaged — and how you average matters enormously.
Consider a document classifier over four classes:
| Class | Support | Precision | Recall | F1 |
|---|---|---|---|---|
| news | 4,000 | 0.95 | 0.97 | 0.96 |
| sport | 3,000 | 0.93 | 0.94 | 0.93 |
| opinion | 800 | 0.71 | 0.62 | 0.66 |
| obituary | 60 | 0.30 | 0.12 | 0.17 |
| Averaging | How | F1 | What it says |
|---|---|---|---|
| Macro | Plain mean of the four F1s | 0.68 | Every class counts equally — the obituary failure is visible |
| Weighted | Mean weighted by support | 0.91 | Dominated by the big classes — the failure disappears |
| Micro | Pool all TP/FP/FN, then compute | 0.92 | Equals accuracy in single-label problems |
Report weighted F1 and the model looks excellent. Report macro F1 and you see it cannot classify obituaries at all. If rare classes matter — and they usually do, since that is often why they were labelled — use macro averaging, and always print the per-class table underneath.
Probabilities that mean something
Every metric so far judges decisions or rankings. If the probability itself feeds a downstream calculation — expected loss, a pricing rule, a triage queue — you need it to be calibrated: among cases assigned 0.30, roughly 30% should be positive.
The Brier score is the mean squared error of the probabilities:
Lower is better, and unlike AUC it punishes confident mistakes rather than just ranking errors. A model can have superb AUC and terrible calibration — ranking everyone correctly while systematically claiming 0.9 for groups whose true rate is 0.6. Random forests and SVMs do this routinely; logistic regression, fitted by maximising likelihood, tends not to.
1from sklearn.calibration import CalibratedClassifierCV, calibration_curve23calibrated = CalibratedClassifierCV(rf, method="isotonic", cv=5)4calibrated.fit(X_train, y_train)56frac_pos, mean_pred = calibration_curve(y_test,7 calibrated.predict_proba(X_test)[:, 1],8 n_bins=10)9# a well-calibrated model has frac_pos ≈ mean_pred in every binChoosing the threshold by cost, not by convention
If you can put a number on each error, the threshold stops being a judgement call. For the screening example: a missed cancer costs an estimated £180,000 in later-stage treatment and outcomes; a false positive costs £400 in follow-up imaging and consultation.
| Threshold | FN | FP | Missed cost | Alarm cost | Total |
|---|---|---|---|---|---|
| 0.90 | 62 | 4 | £11,160,000 | £1,600 | £11,161,600 |
| 0.50 | 28 | 96 | £5,040,000 | £38,400 | £5,078,400 |
| 0.30 | 16 | 310 | £2,880,000 | £124,000 | £3,004,000 |
| 0.10 | 4 | 1,840 | £720,000 | £736,000 | £1,456,000 |
| 0.01 | 0 | 9,920 | £0 | £3,968,000 | £3,968,000 |
The optimum is near 0.10 — far from the default, and the default would have cost three and a half times as much. When costs are genuinely asymmetric, using 0.5 is not neutral; it is an unexamined assumption that both errors cost the same.
1import numpy as np23def cheapest_threshold(y_true, probs, cost_fn, cost_fp):4 best = (np.inf, 0.5)5 for t in np.linspace(0.01, 0.99, 99):6 pred = (probs >= t).astype(int)7 fn = ((y_true == 1) & (pred == 0)).sum()8 fp = ((y_true == 0) & (pred == 1)).sum()9 total = fn * cost_fn + fp * cost_fp10 best = min(best, (total, t))11 return best[1], best[0]Fit the threshold on validation data. Choosing it on the test set turns your final estimate into a tuned number.
A complete evaluation
1from sklearn.metrics import (confusion_matrix, classification_report,2 roc_auc_score, average_precision_score,3 brier_score_loss, balanced_accuracy_score)45probs = model.predict_proba(X_test)[:, 1]6pred = (probs >= chosen_threshold).astype(int)78tn, fp, fn, tp = confusion_matrix(y_test, pred).ravel()9print(f"TP {tp} FP {fp} FN {fn} TN {tn}")10print(classification_report(y_test, pred, digits=3))1112print("balanced accuracy :", round(balanced_accuracy_score(y_test, pred), 3))13print("ROC AUC :", round(roc_auc_score(y_test, probs), 3))14print("average precision :", round(average_precision_score(y_test, probs), 3))15print("positive rate :", round(y_test.mean(), 4), " <- AP baseline")16print("Brier :", round(brier_score_loss(y_test, probs), 4))Printing the positive class rate next to average precision is a small habit that prevents a large mistake. An average precision of 0.19 sounds dreadful until you see the baseline is 0.008, at which point it is a 24-fold improvement over random.
balanced_accuracy_score — the mean of recall on each class — is the honest replacement for accuracy under imbalance. On the do-nothing model that scored 99.2% accuracy, balanced accuracy is exactly 0.50, which is the correct verdict.
Handling imbalance itself
Better metrics reveal the problem; they do not solve it. Three approaches, in order of how often they are the right answer:
| Approach | What it does | Caveat |
|---|---|---|
| Threshold tuning | Move the cutoff on the model you already have | Almost always try this first — it is free |
class_weight="balanced" | Weights the rare class up in the loss | Distorts probabilities; recalibrate if you need them |
| Undersampling the majority | Discards negatives to balance the classes | Throws away real data |
| SMOTE / synthetic oversampling | Generates synthetic minority examples | Must be applied inside cross-validation folds only — oversampling before splitting leaks near-duplicates into validation and inflates every score |
That last caveat is severe enough to be worth stating flatly: applying SMOTE to the full dataset before splitting produces validation scores that are simply wrong, often by 15 percentage points or more, because synthetic points interpolated from a training example land in the validation set.
What this means when you build something
Start every classification evaluation by printing the confusion matrix as raw counts, before any metric. Four numbers tell you what kind of model you have built, and they make it impossible to hide behind an average. If you cannot look at those counts and say which cell your business cares about most, you are not ready to choose a metric.
Then choose deliberately. Under imbalance, report balanced accuracy or macro F1 rather than accuracy, and average precision alongside AUC. Print the per-class breakdown rather than a single averaged figure, because the averaged figure exists to hide exactly the rare-class failure you most need to see.
And treat the threshold as part of the model. Keep the probabilities, sweep the cutoff on validation data with your actual costs attached, and record which threshold you chose and why. A model shipped at 0.5 because that was the default is a model whose most consequential parameter nobody chose.