Machine Learning Foundations

Course Content

Machine Learning Foundations

14 sections · 70 lessons

What is a classification threshold?


What you need to know

A logistic regression, a random forest or a neural network does not really say "fraud" or "not fraud". It says something like "0.83" — a score, often a probability. model.predict() hides a step: it compares that score with a threshold, usually 0.5, and returns 1 if the score is at least that high.

Separate the two jobs

It helps to see a classifier as two parts:

  1. The model ranks — it gives every case a score, ideally higher for positives.
  2. The threshold decides — a business rule turns scores into actions.

Training improves step 1. Step 2 is a policy choice, and it is often where the real value is won or lost.

Using your own threshold in scikit-learn

Python
import numpy as npfrom sklearn.linear_model import LogisticRegressionfrom sklearn.model_selection import train_test_split, TunedThresholdClassifierCVrng = np.random.default_rng(0)n = 30_000y = (rng.random(n) < 0.02).astype(int)              # 2% fraudX = rng.normal(size=(n, 5)) + 1.0 * y[:, None]X_tr, X_te, y_tr, y_te = train_test_split(X, y, stratify=y, random_state=0)model = LogisticRegression().fit(X_tr, y_tr)proba = model.predict_proba(X_te)[:, 1]             # probability of fraudprint("first 5 scores:", proba[:5].round(3))for t in [0.5, 0.1]:    pred = (proba >= t).astype(int)                 # the threshold turns scores into labels    print(f"threshold {t}: flagged {pred.sum()} of {len(pred)}, caught {pred[y_te == 1].sum()} of {y_te.sum()} frauds")tuned = TunedThresholdClassifierCV(LogisticRegression(), scoring="f1").fit(X_tr, y_tr)print("threshold that maximises F1 on CV:", round(tuned.best_threshold_, 2))
Text
first 5 scores: [0.003 0.011 0.001 0.    0.001]threshold 0.5: flagged 56 of 7500, caught 38 of 148 fraudsthreshold 0.1: flagged 319 of 7500, caught 100 of 148 fraudsthreshold that maximises F1 on CV: 0.27

With only 2% fraud, the model's probabilities are mostly small, so 0.5 is a very strict bar: it catches 38 of 148 frauds. At 0.1 it catches 100, at the cost of more flags. Same model, no retraining — only the policy changed.

TunedThresholdClassifierCV (added in scikit-learn 1.5) searches for the threshold that maximises a chosen metric using cross-validation, so the choice is not made on the test set. FixedThresholdClassifier wraps a model with a threshold you pick yourself, so predict() uses it everywhere.

How to choose the threshold

  • By cost — if a missed fraud costs ₹20,000 and a false alarm costs ₹100 of review time, pick the threshold that minimises total cost on validation data.
  • By capacity — if the team can review 300 alerts a day, pick the threshold that produces about 300 flags.
  • By a constraint — "recall at least 90%", then the highest threshold that still meets it.

Choose it on validation data, never the test set, and re-check it after retraining or when the fraud rate changes.

A real-life example

An e-commerce company flags cash-on-delivery orders likely to be refused at the door. Each refused order costs about ₹150 in wasted shipping; asking a customer to prepay instead loses about 20% of flagged orders, and each lost order costs ₹80 in margin. At the default threshold of 0.5, the model flags almost nothing, because only 6% of orders are refused. The analyst computes total cost at every threshold from 0.05 to 0.5 on last month's orders and finds the minimum near 0.18. That threshold goes into a config file, not the model, so the operations team can adjust it during the festive sale without a retrain.

Follow-up questions to expect

  • "Why not always use 0.5?" — Because 0.5 assumes balanced classes and equal costs. With 2% positives, a well-calibrated model rarely gives any case a probability above 0.5, so it flags very little.
  • "Does changing the threshold change the model?" — No. It changes decisions, not the learned scores; ROC-AUC stays the same.
  • "What if the model's probabilities are poorly calibrated?" — The ranking may still be good, so threshold tuning still works, but you cannot read a score of 0.3 as "30% chance". Use CalibratedClassifierCV if probabilities feed into cost calculations.