Course Content
Machine Learning Foundations
14 sections · 70 lessons
Why is accuracy not always a good evaluation metric?
What you need to know
accuracy = correct predictions / all predictionsAccuracy is easy to explain, and it is fine when classes are roughly balanced and both kinds of mistakes cost about the same. It fails in three common situations.
Failure 1: class imbalance
Class imbalance means one class is much rarer than the other. Most interesting problems are imbalanced: fraud, disease, defects, churn. A model can get a high score by always predicting the common class.
1import numpy as np2from sklearn.dummy import DummyClassifier3from sklearn.metrics import accuracy_score, recall_score45rng = np.random.default_rng(0)6y = (rng.random(100_000) < 0.01).astype(int) # 1% of payments are fraud7X = np.zeros((len(y), 1)) # features do not even matter89lazy = DummyClassifier(strategy="most_frequent").fit(X, y) # always says "not fraud"10pred = lazy.predict(X)11print("accuracy:", accuracy_score(y, pred))12print("fraud caught (recall):", recall_score(y, pred))accuracy: 0.99013fraud caught (recall): 0.0A model that has learned nothing scores 99% accuracy. DummyClassifier is a useful habit in its own right: it gives the floor any real model must beat.
Failure 2: all mistakes counted equally
Accuracy adds up a missed fraud and a wrongly blocked customer as "one error each". In reality they cost different amounts: a missed ₹50,000 fraud might cost the full ₹50,000, while a blocked genuine payment costs a support call and some customer goodwill. A metric that ignores this can pick the wrong model.
Failure 3: it hides the probabilities
Accuracy looks only at the final yes/no label, usually made with a 0.5 cut-off. Two models with the same accuracy can have very different probability quality: one might rank fraud cases well and simply need a lower threshold. Threshold-free metrics such as ROC-AUC and PR-AUC show that.
What to use instead
| Situation | Better metric |
|---|---|
| Rare positive class, care about catching it | Recall, precision, F1 for that class, PR-AUC |
| Need to see every type of error | Confusion matrix |
| Threshold not chosen yet | ROC-AUC or PR-AUC |
| Different costs per error | Expected cost in money |
| Classes balanced, equal costs | Accuracy is fine |
A real-life example
A diabetic-retinopathy screening model is tested on 10,000 eye scans from a rural clinic programme, where 3% of patients have the disease. The model reports 97% accuracy. A doctor on the team checks: the model predicts "healthy" for almost everyone and finds only 40 of the 300 sick patients — a recall of 13%. For screening, a missed case can mean preventable blindness, so the team retrains with class weights, lowers the threshold, and reports recall (target above 90%) alongside precision (so the clinic is not flooded with referrals). Accuracy drops to 91%, and the model becomes genuinely useful.
Follow-up questions to expect
- "When is accuracy a fine metric?" — When classes are roughly balanced and false positives and false negatives cost about the same, such as classifying handwritten digits.
- "What is balanced accuracy?" — The average of recall for each class. The always-"not fraud" model gets (1.0 + 0.0) / 2 = 0.5, which exposes it.
- "Would you fix imbalance by resampling?" — Sometimes, but first use the right metric and tune the threshold. Class weights or resampling can help training, but they do not fix a misleading metric.