Machine Learning Foundations

Course Content

Machine Learning Foundations

14 sections · 70 lessons

What is recall, and when is it more important than precision?


What you need to know

Text
recall = TP / (TP + FN)       = positives caught / all actual positives

Recall is also called sensitivity or the true positive rate. It looks only at the cases that are actually positive and asks how many the model found. It ignores false alarms completely — that is precision's job.

The trade-off in code

A common way to push a model toward recall is to tell it the rare class matters more, using class_weight="balanced". Here, 5% of patients have a disease:

Python
import numpy as npfrom sklearn.linear_model import LogisticRegressionfrom sklearn.model_selection import train_test_splitfrom sklearn.metrics import recall_score, precision_scorerng = np.random.default_rng(0)n = 20_000y = (rng.random(n) < 0.05).astype(int)             # 5% of patients have the diseaseX = rng.normal(size=(n, 4)) + 1.2 * y[:, None]     # test results shift for sick patientsX_tr, X_te, y_tr, y_te = train_test_split(X, y, stratify=y, random_state=0)for weight in [None, "balanced"]:    m = LogisticRegression(class_weight=weight).fit(X_tr, y_tr)    pred = m.predict(X_te)    print(f"class_weight={str(weight):8s}  recall={recall_score(y_te, pred):.2f}  "          f"precision={precision_score(y_te, pred):.2f}  flagged={pred.sum()}")
Text
class_weight=None      recall=0.50  precision=0.71  flagged=173class_weight=balanced  recall=0.86  precision=0.28  flagged=767

The default model misses half the sick patients. The balanced model catches 86% of them, but it flags 767 people instead of 173, so precision falls to 0.28. For a screening programme that is the right trade: a false alarm costs one follow-up test, while a miss can cost a life. Lowering the decision threshold has a similar effect (see the Threshold Tuning section).

When recall is the priority

  • Medical screening — a missed cancer or TB case goes untreated.
  • Fraud and security — one missed account takeover can cost lakhs.
  • Safety defects — a cracked brake part reaching a car.
  • Legal and compliance search — missing one relevant document in a court case can be serious.

The two-stage pattern

High recall usually means many false alarms. The standard fix is not to lower recall, but to add a second stage that is more accurate and more expensive: a biopsy after a screening scan, a human analyst after a fraud flag, a second model with richer features. Stage one catches nearly everything; stage two cleans up precision.

Recall alone can be gamed

Predict "positive" for everyone and recall is a perfect 1.0. That is why recall is always reported with precision, or used as "maximise recall subject to precision at least X".

A real-life example

An airport baggage scanner flags bags that might contain a prohibited item. Suppose 1 bag in 10,000 contains one. The system is tuned for very high recall: it flags about 3% of bags, so precision is tiny — almost all flagged bags are harmless. That is acceptable because a flagged bag goes to a human screener who spends one minute on it, while a missed item is a serious security incident. If the airport doubled its passenger count, the team would add screeners or a second automated stage, not cut recall.

Follow-up questions to expect

  • "How do you increase recall?" — Lower the threshold, use class weights or resampling, and add features that reveal the missed positives. Look at the false negatives to see what they have in common.
  • "What is the difference between recall and specificity?" — Recall is the share of actual positives caught; specificity is the share of actual negatives correctly cleared.
  • "Can you have 100% recall?" — Yes, trivially, by flagging everything. The real question is how much precision you must give up to get close to it.