Course Content
Machine Learning Foundations
14 sections · 70 lessons
What does F1-score represent, and why is it useful?
What you need to know
F1 = 2 × precision × recall / (precision + recall)Precision and recall each tell half the story. F1 combines them into one number, which helps when you have to compare many models or pick a single metric to optimise.
Why a harmonic mean and not a normal average?
The ordinary (arithmetic) average is fooled by one very high value. The harmonic mean is dominated by the smaller value, so a model cannot hide one terrible metric behind one excellent one.
1from sklearn.metrics import f1_score, fbeta_score, precision_score, recall_score23actual = [1]*20 + [0]*804lazy = [1] + [0]*19 + [0]*80 # flags one sure case only5useful = [1]*15 + [0]*5 + [1]*10 + [0]*70 # catches 15, 10 false alarms67for name, pred in [("lazy", lazy), ("useful", useful)]:8 p, r = precision_score(actual, pred), recall_score(actual, pred)9 print(f"{name:6s} P={p:.2f} R={r:.2f} mean={(p + r) / 2:.2f} "10 f"F1={f1_score(actual, pred):.2f} F2={fbeta_score(actual, pred, beta=2):.2f}")lazy P=1.00 R=0.05 mean=0.53 F1=0.10 F2=0.06useful P=0.60 R=0.75 mean=0.68 F1=0.67 F2=0.71The lazy model flags one certain case and ignores the rest. Its plain average (0.53) looks respectable. Its F1 (0.10) correctly says it is nearly useless. The useful model has balanced precision and recall, and F1 rewards that.
F-beta: when one side matters more
F1 assumes precision and recall are equally important. F-beta lets you choose:
- F2 (beta = 2) weights recall more — for screening and fraud, where misses hurt more.
- F0.5 (beta = 0.5) weights precision more — for spam filtering or auto-blocking.
In the output above, the useful model's F2 (0.71) is higher than its F1 (0.67) because its recall is higher than its precision.
What F1 cannot see
- True negatives — F1 never uses TN. That is fine when negatives are plentiful and boring, but it means F1 is not suitable when correctly clearing negatives is itself important.
- The threshold — F1 is measured at one threshold. Change the threshold and F1 changes. To compare models regardless of threshold, use PR-AUC.
- Money — F1 does not know that a missed fraud costs ₹40,000 and a false alarm ₹50. If you know the costs, optimise expected cost directly.
Multi-class F1
With more than two classes, F1 is computed per class and averaged. Macro F1 gives each class equal weight, so rare classes count as much as common ones. Weighted F1 weights by class size, so it is dominated by common classes. Micro F1 pools all counts; for single-label multi-class problems, it equals accuracy.
A real-life example
A news app classifies incoming articles into 12 topics, and "Elections" is only 2% of articles but very important during election season. Model A has weighted F1 of 0.91 and Model B 0.89. Looking at per-class F1, Model A scores 0.40 on Elections and Model B scores 0.78. The team uses macro F1 as the main metric, which picks Model B, because they do not want a rare but important topic to be hidden by the common ones.
Follow-up questions to expect
- "When would you not use F1?" — When the costs of false positives and false negatives are very different (use F-beta or expected cost), or when true negatives matter.
- "What is the F1 of a model that predicts everything as positive?" — Recall is 1 and precision equals the positive rate, so with 1% positives, F1 is about 2 × 0.01 / 1.01 ≈ 0.02.
- "Should you pick the threshold that maximises F1?" — Only if precision and recall truly matter equally. Otherwise choose the threshold from business costs or capacity.