Machine Learning Foundations

Course Content

Machine Learning Foundations

14 sections · 70 lessons

When would you prefer ROC–AUC over accuracy or F1-score?


What you need to know

Each metric answers a different question:

MetricQuestion it answersNeeds a threshold?Sensitive to class ratio?
AccuracyWhat share of all decisions was right?YesYes, badly
F1How good are precision and recall together, at this threshold?YesYes
ROC–AUCHow well are positives ranked above negatives?NoNo
PR-AUCHow well are positives found without many false alarms, across thresholds?NoYes

Prefer ROC–AUC when

  • The threshold is not decided yet. Early in a project, you want to know which model ranks best; the threshold comes later from business discussion.
  • Ranking is the product. A sales team calls leads in score order; a reviewer works through a queue from the top.
  • You compare across different class ratios. AUC does not change with the positive rate, so a model tested on a month with 1% fraud and another with 3% fraud can be compared fairly.
  • Classes are balanced or mildly imbalanced, and both error types matter.

Prefer F1 or PR-AUC when

  • Positives are rare and each false alarm costs something. The next example shows why.
  • You have already chosen the operating point — then report F1, or precision and recall, at that point.

Here the same model is scored on data where fraud becomes rarer and rarer:

Python
import numpy as npfrom sklearn.metrics import roc_auc_score, average_precision_scorerng = np.random.default_rng(0)pos = rng.normal(2.0, 1, 1_000)            # scores the model gives to real fraudneg = rng.normal(0.0, 1, 100_000)          # scores it gives to genuine paymentsfor n_neg in [1_000, 10_000, 100_000]:     # same model, rarer and rarer fraud    s = np.r_[pos, neg[:n_neg]]    y = np.r_[np.ones(1_000), np.zeros(n_neg)]    print(f"fraud rate {1_000 / (1_000 + n_neg):5.1%}  ROC-AUC={roc_auc_score(y, s):.3f}  "          f"PR-AUC={average_precision_score(y, s):.3f}")
Text
fraud rate 50.0%  ROC-AUC=0.919  PR-AUC=0.919fraud rate  9.1%  ROC-AUC=0.919  PR-AUC=0.634fraud rate  1.0%  ROC-AUC=0.920  PR-AUC=0.240

ROC–AUC stays at 0.92 in all three cases, because the model's ranking ability did not change. PR-AUC falls from 0.92 to 0.24, because at 1% fraud the same model produces many more false alarms per real catch. Both are right; they answer different questions. If a reviewer has to check every flag, PR-AUC describes their day far better.

Prefer accuracy only when

Classes are roughly balanced and the two errors cost about the same — for example, classifying product photos as "shoe" or "bag".

A real-life example

An ed-tech company builds a model to predict which free-trial users will buy a subscription, and a sales team calls users in score order. About 8% of trial users convert. In the model-selection phase, the team compares gradient boosting, logistic regression and a small neural network by ROC–AUC (0.84, 0.81, 0.83), because the output is a ranked call list and the number of calls per day is not yet fixed. Once the sales head confirms the team can make 400 calls a day, they switch to precision@400 — how many of the top 400 users actually buy — which is the number that decides the sales team's month.

Follow-up questions to expect

  • "Can a model have high AUC and low F1?" — Yes. The ranking can be good while the chosen threshold is poor, or while positives are so rare that even a good ranking gives low precision.
  • "What is partial AUC?" — The area under only part of the ROC curve, such as FPR from 0 to 1%, for when you will only ever operate in that region. scikit-learn supports it through roc_auc_score(..., max_fpr=0.01), which returns a standardised version where 0.5 still means random.
  • "Which metric would you optimise during training?" — Usually a loss like log loss; AUC and F1 are used to select models and thresholds, not as the training objective.