Machine Learning Foundations

Course Content

Machine Learning Foundations

14 sections · 70 lessons

How does class imbalance affect model evaluation?


One model, 1 percent fraud, four very different numbers0.9910.9430.3820.210123accuracyROC-AUCPR-AUCfraudrecall at 0.5Only the minority-class numbers reveal that four frauds in five slip through.
The further a metric sits from the rare class, the better it looks — read the numbers from right to left.

What you need to know

Class imbalance means the classes have very different sizes: 1 fraud per 100 payments, 1 defect per 1,000 parts. The rare class is usually the positive class — the one you are trying to find.

How it distorts each metric

  • Accuracy — dominated by the majority. Predicting "no" always gives 99%.
  • ROC-AUC — uses the false-positive rate, which divides by the number of negatives. With 99,000 negatives, even 1,000 false alarms is only a 1% false-positive rate, so the curve looks excellent while reviewers drown in false alarms.
  • PR-AUC (average precision) — built from precision and recall, both of which focus on the positive class. It drops when the model raises many false alarms, so it is the more honest summary. Its baseline is the positive rate: random guessing on 1% fraud gives PR-AUC of about 0.01, not 0.5.

See it in numbers

Python
import numpy as npfrom sklearn.model_selection import train_test_splitfrom sklearn.linear_model import LogisticRegressionfrom sklearn.metrics import accuracy_score, roc_auc_score, average_precision_score, classification_reportrng = np.random.default_rng(0)n = 50_000y = (rng.random(n) < 0.01).astype(int)               # 1% fraudX = rng.normal(size=(n, 5)) + 1.0 * y[:, None]       # fraud rows shifted a littleX_tr, X_te, y_tr, y_te = train_test_split(X, y, stratify=y, random_state=0)model = LogisticRegression().fit(X_tr, y_tr)proba = model.predict_proba(X_te)[:, 1]print(f"accuracy {accuracy_score(y_te, model.predict(X_te)):.3f}")print(f"ROC-AUC  {roc_auc_score(y_te, proba):.3f}")print(f"PR-AUC   {average_precision_score(y_te, proba):.3f}")print(classification_report(y_te, model.predict(X_te), digits=2, zero_division=0))
Text
accuracy 0.991ROC-AUC  0.943PR-AUC   0.382              precision    recall  f1-score   support           0       0.99      1.00      1.00     12373           1       0.64      0.21      0.32       127    accuracy                           0.99     12500   macro avg       0.82      0.61      0.66     12500weighted avg       0.99      0.99      0.99     12500

Read the numbers in order. Accuracy 0.991 and ROC-AUC 0.943 both sound great. PR-AUC 0.382 is more sober. The report's row for class 1 tells the real story: at the default 0.5 threshold, the model catches only 21% of fraud. The "weighted avg" row is dominated by class 0, so it looks fine too — always read the minority-class row.

What to do about it

  1. Stratify — train_test_split(..., stratify=y) and StratifiedKFold keep the 1% ratio in every split. Without it, a small fold might contain almost no fraud.
  2. Pick minority-focused metrics — precision, recall and F1 for class 1, and PR-AUC.
  3. Read the confusion matrix — see exactly how many frauds were missed.
  4. Tune the threshold — a threshold below 0.5 will raise that 21% recall (see the Threshold Tuning section).
  5. Consider training fixes — class_weight="balanced", resampling, or cost-sensitive learning — then evaluate on the original, untouched class ratio.

A real-life example

A factory uses cameras to spot cracked bottles; 0.2% of bottles are defective. A vendor claims "99.8% accuracy". The quality engineer asks for recall on defects and finds it is 35%: two of every three cracked bottles reach customers. She asks the vendor to report recall and precision on defects, measured on a stratified test set with at least 500 defective examples, so the numbers are not based on a handful of cases. The contract is rewritten around "recall at least 95% with no more than 1 false reject per 200 bottles".

Follow-up questions to expect

  • "Should you balance the test set?" — No. Keep the test set at the real-world ratio, or precision will look better than it will be in production. Resample only the training data, if at all.
  • "Why is PR-AUC better than ROC-AUC here?" — ROC-AUC's false-positive rate is diluted by the huge number of negatives; precision is not, so PR-AUC reflects the false-alarm load a reviewer actually sees.
  • "Does SMOTE fix imbalance?" — SMOTE creates synthetic minority examples for training. It can help some models, but often class weights plus threshold tuning work as well with less risk, and SMOTE must be applied only inside training folds.