Statistics & Math for AI/ML Interviews

Course Content

Statistics & Math for AI/ML Interviews

8 sections · 30 lessons

How is conditional probability mathematically represented?


A checkout funnel is the chain ruleAll visitorsAdd to cart:times 0.30Start checkout,given cart: times 0.50Pay, givencheckout:times 0.8012 percent ofvisitors payEach step's rate is conditioned on every step before it.
Multiplying conditional rates gives the joint probability with no independence assumption at all.

What you need to know

The formula and what each part means

Text
P(A | B) = P(A and B) / P(B)          requires P(B) > 0
  • P(A and B) is the share of all cases where both happen.
  • P(B) is the share of all cases where the condition holds.
  • Dividing rescales the "both" share so that the B-world adds up to 1.

We divide by P(B) because we have thrown away every case where B is false. What remains must be treated as the new whole.

Worked example: an A/B test of a checkout button

10,000 visitors reach checkout. 4,000 are on mobile and 6,000 on desktop. 120 mobile visitors and 300 desktop visitors complete a purchase.

Text
P(mobile)                = 4,000 / 10,000 = 0.40P(purchase and mobile)   =   120 / 10,000 = 0.012P(purchase | mobile)     = 0.012 / 0.40   = 0.03   (3%)P(purchase | desktop)    = 0.030 / 0.60   = 0.05   (5%)

Counting directly gives the same answer: 120 / 4,000 = 3%. The formula and the "shrink the world" count are the same thing.

The multiplication rule and the chain rule

Multiply both sides by P(B):

Text
P(A and B) = P(B) × P(A | B) = P(A) × P(B | A)

Extend it to many events and you get the chain rule:

Text
P(A, B, C) = P(A) × P(B | A) × P(C | A, B)

Worked example: a shopping funnel. 30% of visitors add to cart. Of those, 50% start checkout. Of those, 80% pay.

Text
P(pays) = P(cart) × P(checkout | cart) × P(pay | cart, checkout)        = 0.30 × 0.50 × 0.80 = 0.12

So 12% of visitors pay. Each step's rate is a conditional probability, and a funnel dashboard is just the chain rule drawn as bars.

Language models do exactly this with words:

Text
P("I love cricket") = P("I") × P("love" | "I") × P("cricket" | "I love")

A model like GPT is trained to predict each of those conditional probabilities, one token at a time.

A real-life example

The checkout team's A/B test shows the new button raised overall conversion from 4.2% to 4.4%. Splitting by device with conditional probabilities tells the real story: P(purchase | desktop, new button) stayed at 5%, while P(purchase | mobile, new button) rose from 3.0% to 3.5%. The new button is larger and easier to tap on small screens.

That insight only appears when you condition. Once a significance test confirms the mobile lift is not noise, the team keeps the new button, then spends its next test on the mobile flow, where the gain actually came from.

Follow-up questions to expect

  • "Why must P(B) be greater than zero?" — You cannot shrink the world to an event that never happens; the division would be by zero, so P(A | B) is undefined.
  • "What is the law of total probability?" — P(A) = P(A | B) × P(B) + P(A | not B) × P(not B). Overall conversion is the device-weighted mix of per-device conversions: 0.03 × 0.4 + 0.05 × 0.6 = 0.042.
  • "How does this connect to Bayes' theorem?" — Set the two forms of the multiplication rule equal, P(B) × P(A | B) = P(A) × P(B | A), and divide by P(B).