Course Content
Statistics & Math for AI/ML Interviews
8 sections · 30 lessons
How is conditional probability mathematically represented?
What you need to know
The formula and what each part means
P(A | B) = P(A and B) / P(B) requires P(B) > 0- P(A and B) is the share of all cases where both happen.
- P(B) is the share of all cases where the condition holds.
- Dividing rescales the "both" share so that the B-world adds up to 1.
We divide by P(B) because we have thrown away every case where B is false. What remains must be treated as the new whole.
Worked example: an A/B test of a checkout button
10,000 visitors reach checkout. 4,000 are on mobile and 6,000 on desktop. 120 mobile visitors and 300 desktop visitors complete a purchase.
P(mobile) = 4,000 / 10,000 = 0.40P(purchase and mobile) = 120 / 10,000 = 0.012P(purchase | mobile) = 0.012 / 0.40 = 0.03 (3%)P(purchase | desktop) = 0.030 / 0.60 = 0.05 (5%)Counting directly gives the same answer: 120 / 4,000 = 3%. The formula and the "shrink the world" count are the same thing.
The multiplication rule and the chain rule
Multiply both sides by P(B):
P(A and B) = P(B) × P(A | B) = P(A) × P(B | A)Extend it to many events and you get the chain rule:
P(A, B, C) = P(A) × P(B | A) × P(C | A, B)Worked example: a shopping funnel. 30% of visitors add to cart. Of those, 50% start checkout. Of those, 80% pay.
P(pays) = P(cart) × P(checkout | cart) × P(pay | cart, checkout) = 0.30 × 0.50 × 0.80 = 0.12So 12% of visitors pay. Each step's rate is a conditional probability, and a funnel dashboard is just the chain rule drawn as bars.
Language models do exactly this with words:
P("I love cricket") = P("I") × P("love" | "I") × P("cricket" | "I love")A model like GPT is trained to predict each of those conditional probabilities, one token at a time.
A real-life example
The checkout team's A/B test shows the new button raised overall conversion from 4.2% to 4.4%. Splitting by device with conditional probabilities tells the real story: P(purchase | desktop, new button) stayed at 5%, while P(purchase | mobile, new button) rose from 3.0% to 3.5%. The new button is larger and easier to tap on small screens.
That insight only appears when you condition. Once a significance test confirms the mobile lift is not noise, the team keeps the new button, then spends its next test on the mobile flow, where the gain actually came from.
Follow-up questions to expect
- "Why must P(B) be greater than zero?" — You cannot shrink the world to an event that never happens; the division would be by zero, so P(A | B) is undefined.
- "What is the law of total probability?" — P(A) = P(A | B) × P(B) + P(A | not B) × P(not B). Overall conversion is the device-weighted mix of per-device conversions: 0.03 × 0.4 + 0.05 × 0.6 = 0.042.
- "How does this connect to Bayes' theorem?" — Set the two forms of the multiplication rule equal, P(B) × P(A | B) = P(A) × P(B | A), and divide by P(B).