Course Content
Machine Learning Foundations
14 sections · 70 lessons
When would you prefer ROC–AUC over accuracy or F1-score?
What you need to know
Each metric answers a different question:
| Metric | Question it answers | Needs a threshold? | Sensitive to class ratio? |
|---|---|---|---|
| Accuracy | What share of all decisions was right? | Yes | Yes, badly |
| F1 | How good are precision and recall together, at this threshold? | Yes | Yes |
| ROC–AUC | How well are positives ranked above negatives? | No | No |
| PR-AUC | How well are positives found without many false alarms, across thresholds? | No | Yes |
Prefer ROC–AUC when
- The threshold is not decided yet. Early in a project, you want to know which model ranks best; the threshold comes later from business discussion.
- Ranking is the product. A sales team calls leads in score order; a reviewer works through a queue from the top.
- You compare across different class ratios. AUC does not change with the positive rate, so a model tested on a month with 1% fraud and another with 3% fraud can be compared fairly.
- Classes are balanced or mildly imbalanced, and both error types matter.
Prefer F1 or PR-AUC when
- Positives are rare and each false alarm costs something. The next example shows why.
- You have already chosen the operating point — then report F1, or precision and recall, at that point.
Here the same model is scored on data where fraud becomes rarer and rarer:
1import numpy as np2from sklearn.metrics import roc_auc_score, average_precision_score34rng = np.random.default_rng(0)5pos = rng.normal(2.0, 1, 1_000) # scores the model gives to real fraud6neg = rng.normal(0.0, 1, 100_000) # scores it gives to genuine payments78for n_neg in [1_000, 10_000, 100_000]: # same model, rarer and rarer fraud9 s = np.r_[pos, neg[:n_neg]]10 y = np.r_[np.ones(1_000), np.zeros(n_neg)]11 print(f"fraud rate {1_000 / (1_000 + n_neg):5.1%} ROC-AUC={roc_auc_score(y, s):.3f} "12 f"PR-AUC={average_precision_score(y, s):.3f}")fraud rate 50.0% ROC-AUC=0.919 PR-AUC=0.919fraud rate 9.1% ROC-AUC=0.919 PR-AUC=0.634fraud rate 1.0% ROC-AUC=0.920 PR-AUC=0.240ROC–AUC stays at 0.92 in all three cases, because the model's ranking ability did not change. PR-AUC falls from 0.92 to 0.24, because at 1% fraud the same model produces many more false alarms per real catch. Both are right; they answer different questions. If a reviewer has to check every flag, PR-AUC describes their day far better.
Prefer accuracy only when
Classes are roughly balanced and the two errors cost about the same — for example, classifying product photos as "shoe" or "bag".
A real-life example
An ed-tech company builds a model to predict which free-trial users will buy a subscription, and a sales team calls users in score order. About 8% of trial users convert. In the model-selection phase, the team compares gradient boosting, logistic regression and a small neural network by ROC–AUC (0.84, 0.81, 0.83), because the output is a ranked call list and the number of calls per day is not yet fixed. Once the sales head confirms the team can make 400 calls a day, they switch to precision@400 — how many of the top 400 users actually buy — which is the number that decides the sales team's month.
Follow-up questions to expect
- "Can a model have high AUC and low F1?" — Yes. The ranking can be good while the chosen threshold is poor, or while positives are so rare that even a good ranking gives low precision.
- "What is partial AUC?" — The area under only part of the ROC curve, such as FPR from 0 to 1%, for when you will only ever operate in that region. scikit-learn supports it through
roc_auc_score(..., max_fpr=0.01), which returns a standardised version where 0.5 still means random. - "Which metric would you optimise during training?" — Usually a loss like log loss; AUC and F1 are used to select models and thresholds, not as the training objective.