Machine Learning Foundations

Course Content

Machine Learning Foundations

14 sections · 70 lessons

Why is accuracy not always a good evaluation metric?


A model that always says "not fraud" on 100,000 payments99,01309870Said genuineSaid fraudActual genuineActual fraudAccuracy 99.0 percent; fraud caught: 0 of 987.
When one class is 99 percent of the data, a model that has learned nothing still earns a 99 percent accuracy.

What you need to know

Text
accuracy = correct predictions / all predictions

Accuracy is easy to explain, and it is fine when classes are roughly balanced and both kinds of mistakes cost about the same. It fails in three common situations.

Failure 1: class imbalance

Class imbalance means one class is much rarer than the other. Most interesting problems are imbalanced: fraud, disease, defects, churn. A model can get a high score by always predicting the common class.

Python
import numpy as npfrom sklearn.dummy import DummyClassifierfrom sklearn.metrics import accuracy_score, recall_scorerng = np.random.default_rng(0)y = (rng.random(100_000) < 0.01).astype(int)   # 1% of payments are fraudX = np.zeros((len(y), 1))                       # features do not even matterlazy = DummyClassifier(strategy="most_frequent").fit(X, y)   # always says "not fraud"pred = lazy.predict(X)print("accuracy:", accuracy_score(y, pred))print("fraud caught (recall):", recall_score(y, pred))
Text
accuracy: 0.99013fraud caught (recall): 0.0

A model that has learned nothing scores 99% accuracy. DummyClassifier is a useful habit in its own right: it gives the floor any real model must beat.

Failure 2: all mistakes counted equally

Accuracy adds up a missed fraud and a wrongly blocked customer as "one error each". In reality they cost different amounts: a missed ₹50,000 fraud might cost the full ₹50,000, while a blocked genuine payment costs a support call and some customer goodwill. A metric that ignores this can pick the wrong model.

Failure 3: it hides the probabilities

Accuracy looks only at the final yes/no label, usually made with a 0.5 cut-off. Two models with the same accuracy can have very different probability quality: one might rank fraud cases well and simply need a lower threshold. Threshold-free metrics such as ROC-AUC and PR-AUC show that.

What to use instead

SituationBetter metric
Rare positive class, care about catching itRecall, precision, F1 for that class, PR-AUC
Need to see every type of errorConfusion matrix
Threshold not chosen yetROC-AUC or PR-AUC
Different costs per errorExpected cost in money
Classes balanced, equal costsAccuracy is fine

A real-life example

A diabetic-retinopathy screening model is tested on 10,000 eye scans from a rural clinic programme, where 3% of patients have the disease. The model reports 97% accuracy. A doctor on the team checks: the model predicts "healthy" for almost everyone and finds only 40 of the 300 sick patients — a recall of 13%. For screening, a missed case can mean preventable blindness, so the team retrains with class weights, lowers the threshold, and reports recall (target above 90%) alongside precision (so the clinic is not flooded with referrals). Accuracy drops to 91%, and the model becomes genuinely useful.

Follow-up questions to expect

  • "When is accuracy a fine metric?" — When classes are roughly balanced and false positives and false negatives cost about the same, such as classifying handwritten digits.
  • "What is balanced accuracy?" — The average of recall for each class. The always-"not fraud" model gets (1.0 + 0.0) / 2 = 0.5, which exposes it.
  • "Would you fix imbalance by resampling?" — Sometimes, but first use the right metric and tune the threshold. Class weights or resampling can help training, but they do not fix a misleading metric.