Statistics & Math for AI/ML Interviews

Course Content

Statistics & Math for AI/ML Interviews

8 sections · 30 lessons

What is Bayes’ Theorem, and how would you explain it intuitively in an interview?


100,000 people, a 1-in-1,000 disease, a 99 percent test100,000 people100 have the disease99,900 are healthy99 test positive1 tests negative999 test positive98,901 test negative
The 999 false alarms from the huge healthy group outnumber the 99 real cases, so a positive result means only about a 9 percent chance.

What you need to know

The formula

Text
P(A | B) = P(B | A) × P(A) / P(B)P(B) = P(B | A) × P(A) + P(B | not A) × P(not A)

The second line expands the bottom of the fraction: evidence B can arise either because A is true or because it is not.

It comes straight from the multiplication rule. P(A and B) can be written two ways — P(A | B) × P(B) or P(B | A) × P(A). Set them equal and divide by P(B).

Worked example with counts, not formulas

A disease affects 1 in 1,000 people. The test catches 99% of real cases (sensitivity) and wrongly flags 1% of healthy people (false-positive rate). You test positive. What is the chance you have the disease?

Picture 100,000 people:

Text
Have the disease:   100   → 99 test positive     (true positives)Healthy:         99,900   → 999 test positive    (false positives)All positives: 99 + 999 = 1,098P(disease | positive) = 99 / 1,098 ≈ 9%

Only 9%. The test is good, but there are so many healthy people that 1% of them (999) outnumbers all the sick people who tested positive (99). Counting people like this — natural frequencies — is the easiest way to explain Bayes in an interview.

Python
def posterior(prior, sensitivity, false_positive_rate):    """P(disease | positive test) by Bayes' theorem."""    evidence = sensitivity * prior + false_positive_rate * (1 - prior)    return sensitivity * prior / evidenceprint(round(posterior(0.001, 0.99, 0.01), 3))   # rare diseaseprint(round(posterior(0.10, 0.99, 0.01), 3))    # common disease, same test
Text
0.090.917

The same test gives 9% or 92% depending only on the prior. That one comparison shows why base rates matter.

A real-life example

A bank's fraud model flags UPI payments. It catches 95% of fraud and flags 2% of genuine payments. Fraud is 1 in 2,000 payments. Out of 1,000,000 payments:

Text
Fraud:       500   → 475 flaggedGenuine: 999,500   → 19,990 flaggedP(fraud | flagged) = 475 / 20,465 ≈ 2.3%

About 43 of every 44 flagged payments are genuine. If every flag blocks the payment, thousands of customers get angry every day. The team uses Bayes to set policy: flags trigger a cheap step-up check, such as an OTP, not a block, and they work on reducing the false-positive rate, which matters far more than raising recall.

An everyday version from cricket. You hear on the radio that a delivery "turned sharply" but missed who bowled it. The spinner bowls 20% of the overs and turns the ball sharply on 60% of deliveries. The seamers bowl 80% of the overs and get sharp turn on 5%. P(spinner | sharp turn) = 0.2 × 0.6 / (0.2 × 0.6 + 0.8 × 0.05) = 0.12 / 0.16 = 75%. Strong evidence, but a one-in-four chance it was a seamer remains, because seamers bowl so much more.

Follow-up questions to expect

  • "How do you improve P(disease | positive)?" — Lower the false-positive rate, or test a group where the disease is more common (a higher prior), or run a second independent test on the positives.
  • "What is P(B) called and why is it needed?" — The evidence or marginal likelihood; it rescales the result so the posteriors over all hypotheses add to 1.
  • "How does this relate to precision?" — Precision is exactly P(actually positive | predicted positive), so it falls when the positive class is rare, even with a good model.