Statistics & Math for AI/ML Interviews

Course Content

Statistics & Math for AI/ML Interviews

8 sections · 30 lessons

How does Bayes’ Theorem help update beliefs when new data arrives?


Conversion rate of a checkout button, day by dayPrior: every rate from 0 to 1 equally likelyDay 1, 30 of 1,000 buy: 3.1percent, range 2.1 to 4.3Day 2, 45 of 1,000 buy: 3.8percent, range 3.0 to 4.7Day 3, 38 of 1,000 buy: 3.8percent, range 3.1 to 4.5
Each day's posterior becomes the next day's prior, so the estimate settles and its range narrows as evidence piles up.

What you need to know

Text
posterior is proportional to prior × likelihoodtoday's posterior  →  tomorrow's prior

"Proportional to" means you can skip P(evidence) while updating and rescale at the end so everything sums to 1.

A concrete update: conversion rate of a checkout button

For a rate like "fraction of visitors who buy", a convenient prior is the Beta distribution, Beta(a, b), which you can read as "a - 1 past successes and b - 1 past failures". Start with Beta(1, 1): every rate from 0 to 1 is equally plausible. After s buys and f non-buys, the posterior is Beta(1 + s, 1 + f). The update is just adding counts.

Python
from scipy.stats import betaa, b = 1, 1                                  # flat prior: any rate is possiblefor day, (buys, visits) in enumerate([(30, 1000), (45, 1000), (38, 1000)], 1):    a, b = a + buys, b + (visits - buys)     # yesterday's posterior is today's prior    low, high = beta.ppf([0.025, 0.975], a, b)    print(f"day {day}: mean={a/(a+b):.4f}  95% interval=({low:.4f}, {high:.4f})")
Text
day 1: mean=0.0309  95% interval=(0.0211, 0.0425)day 2: mean=0.0380  95% interval=(0.0300, 0.0468)day 3: mean=0.0380  95% interval=(0.0314, 0.0451)

Watch the interval. The best guess settles near 3.8%, and the range of plausible values narrows each day as evidence accumulates. Feeding all 3,000 visits at once gives exactly the same Beta(114, 2888).

A real-life example

An e-commerce team tests a new checkout button. After a week: old button 300 purchases from 10,000 visitors; new button 345 from 10,000. They put a Beta(1, 1) prior on each rate and sample from both posteriors:

Python
import numpy as nprng = np.random.default_rng(0)old = rng.beta(1 + 300, 1 + 9_700, 200_000)   # old button: 300 buys / 10,000new = rng.beta(1 + 345, 1 + 9_655, 200_000)   # new button: 345 buys / 10,000print("P(new button is better):", (new > old).mean().round(3))
Text
P(new button is better): 0.964

The team can tell the product manager, "there is about a 96% chance the new button converts better". That is easier to act on than a p-value. They can also keep updating daily and stop when the probability is high enough — though stopping rules still need to be agreed in advance, because peeking at any test many times raises the chance of acting on noise.

The same pattern runs in Thompson sampling for recommendations: each item has a Beta posterior on its click rate, the system samples from each and shows the winner, so it explores uncertain items and exploits proven ones automatically.

Follow-up questions to expect

  • "Does the order of the data matter?" — No, for independent observations the final posterior is the same whatever the order or batching.
  • "What if the world changes over time?" — Old evidence becomes misleading, so teams discount it (forgetting factors) or use a sliding window so the prior does not become overconfident.
  • "Bayesian or frequentist A/B testing?" — Bayesian gives a direct probability that B beats A and handles prior knowledge; frequentist gives p-values with well-known error guarantees. Both need a pre-agreed stopping rule.