Machine Learning Foundations

Course Content

Machine Learning Foundations

14 sections · 70 lessons

How does changing the threshold affect precision and recall?


Sliding the threshold on one fraud model51.000.03560.680.263190.310.681,6540.090.96flaggedprecisionrecallthreshold 0.9threshold 0.5threshold 0.1threshold 0.01Highlighted: the row that fits a team able to review about 300 alerts.
No row is the correct one — each threshold is a business policy that trades false alarms for catches.

What you need to know

Picture every case sorted by its score, highest first. The threshold is a line drawn across that list. Everything above the line is flagged.

  • Move the line down (lower threshold) → more cases flagged → more real positives caught (recall up) → but more negatives flagged too (precision usually down).
  • Move the line up (higher threshold) → fewer, more confident flags → precision up, recall down.

Recall can only go up or stay the same as the threshold falls, because you only ever add flags. Precision usually falls, but not always — if the next few cases you add are mostly positives, precision can briefly rise.

The trade-off in numbers

Python
import numpy as npfrom sklearn.linear_model import LogisticRegressionfrom sklearn.model_selection import train_test_splitfrom sklearn.metrics import precision_score, recall_scorerng = np.random.default_rng(0)n = 30_000y = (rng.random(n) < 0.02).astype(int)              # 2% fraudX = rng.normal(size=(n, 5)) + 1.0 * y[:, None]X_tr, X_te, y_tr, y_te = train_test_split(X, y, stratify=y, random_state=0)proba = LogisticRegression().fit(X_tr, y_tr).predict_proba(X_te)[:, 1]print("threshold  flagged  precision  recall")for t in [0.9, 0.5, 0.3, 0.1, 0.05, 0.01]:    pred = (proba >= t).astype(int)    print(f"{t:>9}  {pred.sum():>7}  {precision_score(y_te, pred, zero_division=0):>9.2f}"          f"  {recall_score(y_te, pred):>6.2f}")
Text
threshold  flagged  precision  recall      0.9        5       1.00    0.03      0.5       56       0.68    0.26      0.3      104       0.52    0.36      0.1      319       0.31    0.68     0.05      586       0.19    0.76     0.01     1654       0.09    0.96

Read it from top to bottom. At 0.9, every flag is right but the model catches 3% of fraud. At 0.01, it catches 96% of fraud, but only 9% of its 1,654 flags are real. No row is "correct". Each is a different business policy.

The extremes

  • Threshold at or below the lowest score — everything is flagged. Recall = 1. Precision = the positive rate (2% here).
  • Threshold above the highest score — nothing is flagged. Recall = 0. Precision is 0 / 0, undefined.

The precision–recall curve

sklearn.metrics.precision_recall_curve computes precision and recall at every possible threshold. Plotted with recall on the x-axis and precision on the y-axis, it usually slopes down to the right. A better model's curve sits higher and further right. The area under it is PR-AUC (average precision), a threshold-free summary.

Picking the operating point

The operating point is the threshold you actually deploy. Common rules:

RuleExample
Meet a recall target"Catch at least 75% of fraud" → 0.05 in the table
Meet a precision floor"At least half of alerts must be real" → 0.3
Fit capacity"We can review about 300 alerts per 7,500 payments" → 0.1
Minimise costCompute ₹ cost of FN and FP at each threshold, pick the lowest

A real-life example

A hospital's sepsis early-warning model sends an alert to the nurse on duty. At a threshold of 0.2, it catches 85% of sepsis cases, but nurses get 30 alerts per shift and most are false. Within weeks, nurses start dismissing alerts without looking — alert fatigue — and the effective recall collapses. The clinical team raises the threshold to 0.45, which drops recall to 70% but cuts alerts to 8 per shift, each one taken seriously. They add a second, lower "watch list" tier for scores between 0.2 and 0.45 that is reviewed once per shift instead of paging the nurse. Two thresholds, two actions.

Follow-up questions to expect

  • "Can precision go up when you lower the threshold?" — Occasionally, for a short stretch, if the newly flagged cases are mostly true positives. Overall, the trend is downward.
  • "How would you set a threshold for 90% recall?" — Use precision_recall_curve on validation data, find the highest threshold with recall at least 0.90, and check the precision there is acceptable.
  • "Can you use different thresholds for different groups?" — Technically yes, such as a lower threshold for high-value transactions. It must be justified by cost and checked for fairness and legal rules.