Statistics & Math for AI/ML Interviews

Course Content

Statistics & Math for AI/ML Interviews

8 sections · 30 lessons

What is conditional probability, and how is it used in classification models?


1,000 emails, conditioned on the word free15050200150650800spamnot spamtotalcontains freeno freeKeep only the shaded row: 150 of 200 is 0.75.
Conditioning throws away every row where the condition is false, and the spam rate jumps from 0.30 to 0.75.

What you need to know

The intuition: shrink the world

Out of 1,000 emails, 300 are spam. So without knowing anything else, P(spam) = 300 / 1,000 = 0.30. Now you are told the email contains the word "free". Ignore every email without "free" — that is the shrinking. Suppose 200 emails contain "free", and 150 of those are spam. Inside this smaller world, P(spam | "free") = 150 / 200 = 0.75.

SpamNot spamTotal
Contains "free"15050200
No "free"150650800
Total3007001,000

The word "free" raised the probability of spam from 0.30 to 0.75. Emails without it drop to 150 / 800 ≈ 0.19.

The same table in pandas. normalize="index" divides each row by its row total, which is exactly "condition on the row":

Python
import pandas as pdemails = pd.DataFrame({    "has_free": [True] * 200 + [False] * 800,    "is_spam":  [True] * 150 + [False] * 50 + [True] * 150 + [False] * 650,})print("P(spam) =", emails["is_spam"].mean())print(pd.crosstab(emails["has_free"], emails["is_spam"], normalize="index"))
Text
P(spam) = 0.3is_spam    False   Truehas_freeFalse     0.8125  0.1875True      0.2500  0.7500

How classifiers use it

Every probabilistic classifier estimates P(class | features) for a new input:

  • Logistic regression directly outputs P(y = 1 | x).
  • Naive Bayes builds P(class | features) from P(features | class) using Bayes' theorem.
  • Neural networks with softmax output P(class | input) for each class.

The decision is a separate step. You pick a threshold based on costs. A spam filter might require P(spam | email) above 0.9 before hiding an email, because hiding a real email is expensive. A fraud screen might send anything above 0.3 to manual review, because missing fraud is expensive.

Calibration checks whether these conditional probabilities are honest: of all emails the model scored around 0.8, are about 80% really spam?

A real-life example

A telecom company's churn model outputs P(churn | features) for each customer. Customer A: 0.72. Customer B: 0.08. Retention offers cost ₹300 each and a lost customer costs about ₹4,000 in future revenue. Assuming an offer keeps the customer, the expected loss from not acting on customer A is 0.72 × 4,000 = ₹2,880, far more than the ₹300 offer. For customer B it is 0.08 × 4,000 = ₹320, roughly break-even.

So the team sends offers to customers with P(churn | features) above about 0.075, which is 300 / 4,000. The threshold comes straight from the conditional probability and the business costs, not from the default 0.5.

Follow-up questions to expect

  • "Is P(A | B) the same as P(B | A)?" — No. P(spam | "free") = 0.75, but P("free" | spam) = 150 / 300 = 0.5. Bayes' theorem converts one into the other.
  • "Why not always use a 0.5 threshold?" — 0.5 is only best when both errors cost the same and the classes are balanced; choose the threshold from the cost of false positives versus false negatives.
  • "What is a calibrated model?" — One where predicted probabilities match observed frequencies; check with a reliability diagram and fix with Platt scaling, isotonic regression or temperature scaling.