Course Content
Statistics & Math for AI/ML Interviews
8 sections · 30 lessons
What is conditional probability, and how is it used in classification models?
What you need to know
The intuition: shrink the world
Out of 1,000 emails, 300 are spam. So without knowing anything else, P(spam) = 300 / 1,000 = 0.30. Now you are told the email contains the word "free". Ignore every email without "free" — that is the shrinking. Suppose 200 emails contain "free", and 150 of those are spam. Inside this smaller world, P(spam | "free") = 150 / 200 = 0.75.
| Spam | Not spam | Total | |
|---|---|---|---|
| Contains "free" | 150 | 50 | 200 |
| No "free" | 150 | 650 | 800 |
| Total | 300 | 700 | 1,000 |
The word "free" raised the probability of spam from 0.30 to 0.75. Emails without it drop to 150 / 800 ≈ 0.19.
The same table in pandas. normalize="index" divides each row by its row total, which is exactly "condition on the row":
1import pandas as pd23emails = pd.DataFrame({4 "has_free": [True] * 200 + [False] * 800,5 "is_spam": [True] * 150 + [False] * 50 + [True] * 150 + [False] * 650,6})7print("P(spam) =", emails["is_spam"].mean())8print(pd.crosstab(emails["has_free"], emails["is_spam"], normalize="index"))P(spam) = 0.3is_spam False Truehas_freeFalse 0.8125 0.1875True 0.2500 0.7500How classifiers use it
Every probabilistic classifier estimates P(class | features) for a new input:
- Logistic regression directly outputs P(y = 1 | x).
- Naive Bayes builds P(class | features) from P(features | class) using Bayes' theorem.
- Neural networks with softmax output P(class | input) for each class.
The decision is a separate step. You pick a threshold based on costs. A spam filter might require P(spam | email) above 0.9 before hiding an email, because hiding a real email is expensive. A fraud screen might send anything above 0.3 to manual review, because missing fraud is expensive.
Calibration checks whether these conditional probabilities are honest: of all emails the model scored around 0.8, are about 80% really spam?
A real-life example
A telecom company's churn model outputs P(churn | features) for each customer. Customer A: 0.72. Customer B: 0.08. Retention offers cost ₹300 each and a lost customer costs about ₹4,000 in future revenue. Assuming an offer keeps the customer, the expected loss from not acting on customer A is 0.72 × 4,000 = ₹2,880, far more than the ₹300 offer. For customer B it is 0.08 × 4,000 = ₹320, roughly break-even.
So the team sends offers to customers with P(churn | features) above about 0.075, which is 300 / 4,000. The threshold comes straight from the conditional probability and the business costs, not from the default 0.5.
Follow-up questions to expect
- "Is P(A | B) the same as P(B | A)?" — No. P(spam | "free") = 0.75, but P("free" | spam) = 150 / 300 = 0.5. Bayes' theorem converts one into the other.
- "Why not always use a 0.5 threshold?" — 0.5 is only best when both errors cost the same and the classes are balanced; choose the threshold from the cost of false positives versus false negatives.
- "What is a calibrated model?" — One where predicted probabilities match observed frequencies; check with a reliability diagram and fix with Platt scaling, isotonic regression or temperature scaling.