Machine Learning Foundations

Course Content

Machine Learning Foundations

14 sections · 70 lessons

What is precision, and when is it more important than accuracy?


What you need to know

Text
precision = TP / (TP + FP)          = correct flags / all flags

Precision looks only at the cases the model flagged. It ignores everything the model left alone, so it says nothing about how many positives were missed — that is recall's job.

Precision and recall pull in opposite directions

A model can get high precision by being cautious and flagging only the obvious cases. It pays for that in recall:

Python
from sklearn.metrics import precision_score, recall_scoreactual  = [1, 1, 1, 1, 0, 0, 0, 0, 0, 0]      # 1 = spamcautious = [1, 1, 0, 0, 0, 0, 0, 0, 0, 0]     # flags only obvious spameager    = [1, 1, 1, 1, 1, 1, 1, 0, 0, 0]     # flags anything suspiciousfor name, pred in [("cautious", cautious), ("eager", eager)]:    print(f"{name:8s} precision={precision_score(actual, pred):.2f}  recall={recall_score(actual, pred):.2f}")
Text
cautious precision=1.00  recall=0.50eager    precision=0.57  recall=1.00

The cautious filter is never wrong when it flags spam, but it lets half of the spam through. The eager filter catches all spam but sends 3 real emails to junk. Which is better depends on what a false positive costs.

When precision beats accuracy as the target

  • Spam filtering — a real bank statement or job offer in the junk folder is worse than one extra spam in the inbox.
  • Blocking payments or suspending accounts — each false positive is a genuine customer who cannot pay for groceries.
  • Human review queues — if a team of 5 can check 500 alerts a day, low precision wastes their time on false alarms, and after weeks of false alarms people start ignoring alerts altogether (alert fatigue).
  • Recommendations — showing 10 products where only 2 are relevant makes the feature look broken.

Accuracy fails in all of these because negatives dominate. A spam filter that never flags anything can still be 90% accurate if 10% of email is spam.

Precision depends on how common positives are

This surprises many candidates. Keep the model the same but make positives rarer, and precision falls, because the same false-positive rate now applies to many more negatives. A fraud model with 1% false-positive rate has decent precision when 10% of payments are fraud, and poor precision when 0.1% are. Always quote precision together with the base rate of the data it was measured on.

A real-life example

A bank's card team builds a model that sends an SMS "Did you make this purchase?" and temporarily blocks the card. At the first threshold, it flags 2,000 transactions a day with a precision of 5% — 1,900 genuine customers get their card blocked daily, and complaints to the call centre triple. The team raises the threshold until precision reaches 40% (flagging about 300 a day), and sends lower-risk cases to a softer check (a one-time password) instead of a block. They accept catching somewhat less fraud with the hard block, because a genuine customer who is blocked twice often closes the card.

Follow-up questions to expect

  • "What is precision when the model predicts no positives at all?" — 0 / 0, which is undefined. scikit-learn returns 0 with a warning, controlled by the zero_division argument.
  • "How do you increase precision?" — Raise the decision threshold, add features that separate the confusing negatives, or add a second-stage check on flagged cases.
  • "What is precision@k?" — Precision among the top k ranked predictions, such as the top 10 recommendations or the 500 alerts a team can review.