Course Content
Machine Learning Foundations
14 sections · 70 lessons
Why is ROC–AUC considered a threshold-independent metric?
What you need to know
Most metrics — accuracy, precision, recall, F1 — need hard labels, so they depend on the threshold you picked. ROC–AUC (area under the Receiver Operating Characteristic curve) instead looks at the scores themselves.
Only the order of scores matters
If you change the scores in any way that keeps their order — multiply them, cube them, take a log — AUC does not change. Threshold-based metrics do:
1import numpy as np2from sklearn.linear_model import LogisticRegression3from sklearn.model_selection import train_test_split4from sklearn.metrics import roc_auc_score56rng = np.random.default_rng(0)7n = 30_0008y = (rng.random(n) < 0.02).astype(int)9X = rng.normal(size=(n, 5)) + 1.0 * y[:, None]10X_tr, X_te, y_tr, y_te = train_test_split(X, y, stratify=y, random_state=0)11proba = LogisticRegression().fit(X_tr, y_tr).predict_proba(X_te)[:, 1]1213squashed = proba ** 3 # same order, very different values14for name, s in [("original", proba), ("cubed", squashed)]:15 print(f"{name:8s} AUC={roc_auc_score(y_te, s):.3f} "16 f"recall@0.5={((s >= 0.5) & (y_te == 1)).sum() / y_te.sum():.2f}")1718pos, neg = proba[y_te == 1], proba[y_te == 0]19wins = (pos[:, None] > neg[None, :]).mean() # every fraud vs every genuine payment20print(f"share of (fraud, genuine) pairs ranked correctly = {wins:.3f}")original AUC=0.954 recall@0.5=0.26cubed AUC=0.954 recall@0.5=0.10share of (fraud, genuine) pairs ranked correctly = 0.954Cubing the scores changes recall at 0.5 from 0.26 to 0.10, but AUC stays at 0.954. The last line checks the definition directly: it compares every fraud score with every genuine score, and the share of pairs where fraud scored higher is exactly the AUC. That pair view is the easiest way to explain AUC in an interview.
How to read AUC values
- 0.5 — no better than random ordering.
- 0.7–0.8 — useful but weak separation.
- 0.9 and above — strong ranking. If you see 0.99+ on a hard problem, suspect leakage first.
- Below 0.5 — worse than random; usually labels or score direction are flipped.
What AUC does not tell you
- Precision at your threshold. AUC treats a false positive rate of 1% as small, but 1% of 1 crore payments is 1 lakh false alarms.
- Calibration. Cubed scores have the same AUC but are no longer valid probabilities.
- Performance where it matters. Two models can have the same AUC while one is better in the low-false-positive region you actually use.
So AUC is a good comparison metric before choosing a threshold, and it should be paired with precision and recall at the operating point.
A real-life example
A credit-card issuer is choosing between two vendors' fraud scores. Vendor A scores between 0 and 1; Vendor B outputs a "risk score" between 0 and 999. Their accuracy numbers cannot be compared, because each was measured at a different threshold on a different scale. The analyst computes ROC–AUC for both on the same month of labelled transactions: A gets 0.93, B gets 0.95. Because AUC depends only on ranking, the different scales do not matter. She then checks recall at the false positive rate the bank can tolerate (0.5%) and confirms B still wins there before signing the contract.
Follow-up questions to expect
- "What does an AUC of 0.5 mean?" — The model ranks a random positive above a random negative only half the time, which is no better than a coin toss.
- "Is AUC affected by class imbalance?" — AUC does not change with the positive rate, because TPR and FPR are each computed within one class. That is exactly why it can hide a heavy false-alarm load when positives are rare; PR-AUC does change with the rate and shows it.
- "Can you compute AUC for multi-class problems?" — Yes, with one-vs-rest or one-vs-one averaging:
roc_auc_score(y, proba, multi_class="ovr").