Course Content
Machine Learning Foundations
14 sections · 70 lessons
What is precision, and when is it more important than accuracy?
What you need to know
precision = TP / (TP + FP) = correct flags / all flagsPrecision looks only at the cases the model flagged. It ignores everything the model left alone, so it says nothing about how many positives were missed — that is recall's job.
Precision and recall pull in opposite directions
A model can get high precision by being cautious and flagging only the obvious cases. It pays for that in recall:
1from sklearn.metrics import precision_score, recall_score23actual = [1, 1, 1, 1, 0, 0, 0, 0, 0, 0] # 1 = spam4cautious = [1, 1, 0, 0, 0, 0, 0, 0, 0, 0] # flags only obvious spam5eager = [1, 1, 1, 1, 1, 1, 1, 0, 0, 0] # flags anything suspicious67for name, pred in [("cautious", cautious), ("eager", eager)]:8 print(f"{name:8s} precision={precision_score(actual, pred):.2f} recall={recall_score(actual, pred):.2f}")cautious precision=1.00 recall=0.50eager precision=0.57 recall=1.00The cautious filter is never wrong when it flags spam, but it lets half of the spam through. The eager filter catches all spam but sends 3 real emails to junk. Which is better depends on what a false positive costs.
When precision beats accuracy as the target
- Spam filtering — a real bank statement or job offer in the junk folder is worse than one extra spam in the inbox.
- Blocking payments or suspending accounts — each false positive is a genuine customer who cannot pay for groceries.
- Human review queues — if a team of 5 can check 500 alerts a day, low precision wastes their time on false alarms, and after weeks of false alarms people start ignoring alerts altogether (alert fatigue).
- Recommendations — showing 10 products where only 2 are relevant makes the feature look broken.
Accuracy fails in all of these because negatives dominate. A spam filter that never flags anything can still be 90% accurate if 10% of email is spam.
Precision depends on how common positives are
This surprises many candidates. Keep the model the same but make positives rarer, and precision falls, because the same false-positive rate now applies to many more negatives. A fraud model with 1% false-positive rate has decent precision when 10% of payments are fraud, and poor precision when 0.1% are. Always quote precision together with the base rate of the data it was measured on.
A real-life example
A bank's card team builds a model that sends an SMS "Did you make this purchase?" and temporarily blocks the card. At the first threshold, it flags 2,000 transactions a day with a precision of 5% — 1,900 genuine customers get their card blocked daily, and complaints to the call centre triple. The team raises the threshold until precision reaches 40% (flagging about 300 a day), and sends lower-risk cases to a softer check (a one-time password) instead of a block. They accept catching somewhat less fraud with the hard block, because a genuine customer who is blocked twice often closes the card.
Follow-up questions to expect
- "What is precision when the model predicts no positives at all?" — 0 / 0, which is undefined. scikit-learn returns 0 with a warning, controlled by the
zero_divisionargument. - "How do you increase precision?" — Raise the decision threshold, add features that separate the confusing negatives, or add a second-stage check on flagged cases.
- "What is precision@k?" — Precision among the top k ranked predictions, such as the top 10 recommendations or the 500 alerts a team can review.