Course Content
Machine Learning Foundations
14 sections · 70 lessons
What factors decide which evaluation metric to use?
What you need to know
A metric is the number you use to judge a model. Choosing it is a design decision, because the model you pick — and the threshold you set — follows from it. Six questions drive the choice.
1. What does each mistake cost?
- False negative expensive (missed fraud, missed tumour) → optimise recall.
- False positive expensive (wrongly blocked payment, real email sent to spam) → optimise precision.
- Both matter → F1, or better, an expected-cost number in rupees.
2. How balanced are the classes?
Roughly balanced → accuracy is acceptable. Heavily imbalanced (1% or less) → precision, recall, F1 and PR-AUC, which focuses on the rare class.
3. Ranking or decision?
If the output is a ranked list — which leads a sales team calls first, which transactions a reviewer checks — use a ranking metric such as ROC-AUC or precision@k (how many of the top k are truly positive). If the output is an automatic yes/no, use threshold-based metrics at your chosen threshold.
4. Do the probabilities themselves matter?
If downstream code multiplies the probability by a rupee amount ("expected loss = 0.3 × ₹40,000"), the probabilities must be calibrated — a predicted 30% should come true about 30% of the time. Use log loss or the Brier score, which punish over-confident wrong probabilities.
5. Regression: how much do big errors hurt?
For predicting a number, the two common metrics behave differently:
1import numpy as np2from sklearn.metrics import mean_absolute_error, root_mean_squared_error34actual = np.array([30, 25, 40, 35, 28, 32, 45, 30, 27, 33]) # delivery minutes5model_a = actual + np.array([3, -3, 3, -3, 3, -3, 3, -3, 3, -3]) # always off by 36model_b = actual + np.array([0, 0, 0, 0, 0, 0, 0, 0, 0, 30]) # perfect, except one 30-min miss78for name, pred in [("A", model_a), ("B", model_b)]:9 print(name, "MAE", mean_absolute_error(actual, pred),10 "RMSE", round(root_mean_squared_error(actual, pred), 1))A MAE 3.0 RMSE 3.0B MAE 3.0 RMSE 9.5MAE (mean absolute error) says the two models are equally good. RMSE (root mean squared error) squares each error before averaging, so model B's single 30-minute miss dominates. If one very late delivery makes a customer cancel their subscription, RMSE matches the business better. If every minute costs the same, MAE is fairer and easier to explain ("off by 3 minutes on average").
6. Can stakeholders understand it?
A metric nobody understands will not drive decisions. "We catch 82% of fraud and 1 in 5 alerts is real" beats "F1 = 0.31".
A real-life example
An e-commerce company builds three models, and each needs a different metric:
| Model | Main metric | Why |
|---|---|---|
| Product recommendations on the home page | Precision@10 | Only 10 slots are shown; what matters is how many are relevant |
| Cash-on-delivery fraud check | Recall at a fixed precision | Missed fraud costs money, but too many blocks annoy genuine buyers |
| Delivery-date promise | MAE in days, plus % of orders later than promised | Customers remember late deliveries more than early ones |
Using ROC-AUC for all three would be "correct" and nearly useless for decisions.
Follow-up questions to expect
- "Why not just optimise accuracy and adjust later?" — The metric decides which model and threshold you pick. Optimising the wrong one can select a model that is worse on the thing you care about.
- "MAE or RMSE for house prices?" — RMSE punishes large errors more; MAE is more robust to a few extreme houses. Many teams also report percentage error, because ₹5 lakh off matters more on a ₹30 lakh flat than on a ₹3 crore villa.
- "What is a business metric versus a model metric?" — A model metric (recall) is measured offline on labels; a business metric (fraud loss in rupees, conversion rate) is what the company tracks. A good model metric moves the business metric.