Course Content
Statistics & Math for AI/ML Interviews
8 sections · 30 lessons
How does Bayes’ Theorem help update beliefs when new data arrives?
What you need to know
posterior is proportional to prior × likelihoodtoday's posterior → tomorrow's prior"Proportional to" means you can skip P(evidence) while updating and rescale at the end so everything sums to 1.
A concrete update: conversion rate of a checkout button
For a rate like "fraction of visitors who buy", a convenient prior is the Beta distribution, Beta(a, b), which you can read as "a - 1 past successes and b - 1 past failures". Start with Beta(1, 1): every rate from 0 to 1 is equally plausible. After s buys and f non-buys, the posterior is Beta(1 + s, 1 + f). The update is just adding counts.
1from scipy.stats import beta23a, b = 1, 1 # flat prior: any rate is possible4for day, (buys, visits) in enumerate([(30, 1000), (45, 1000), (38, 1000)], 1):5 a, b = a + buys, b + (visits - buys) # yesterday's posterior is today's prior6 low, high = beta.ppf([0.025, 0.975], a, b)7 print(f"day {day}: mean={a/(a+b):.4f} 95% interval=({low:.4f}, {high:.4f})")day 1: mean=0.0309 95% interval=(0.0211, 0.0425)day 2: mean=0.0380 95% interval=(0.0300, 0.0468)day 3: mean=0.0380 95% interval=(0.0314, 0.0451)Watch the interval. The best guess settles near 3.8%, and the range of plausible values narrows each day as evidence accumulates. Feeding all 3,000 visits at once gives exactly the same Beta(114, 2888).
A real-life example
An e-commerce team tests a new checkout button. After a week: old button 300 purchases from 10,000 visitors; new button 345 from 10,000. They put a Beta(1, 1) prior on each rate and sample from both posteriors:
1import numpy as np23rng = np.random.default_rng(0)4old = rng.beta(1 + 300, 1 + 9_700, 200_000) # old button: 300 buys / 10,0005new = rng.beta(1 + 345, 1 + 9_655, 200_000) # new button: 345 buys / 10,0006print("P(new button is better):", (new > old).mean().round(3))P(new button is better): 0.964The team can tell the product manager, "there is about a 96% chance the new button converts better". That is easier to act on than a p-value. They can also keep updating daily and stop when the probability is high enough — though stopping rules still need to be agreed in advance, because peeking at any test many times raises the chance of acting on noise.
The same pattern runs in Thompson sampling for recommendations: each item has a Beta posterior on its click rate, the system samples from each and shows the winner, so it explores uncertain items and exploits proven ones automatically.
Follow-up questions to expect
- "Does the order of the data matter?" — No, for independent observations the final posterior is the same whatever the order or batching.
- "What if the world changes over time?" — Old evidence becomes misleading, so teams discount it (forgetting factors) or use a sliding window so the prior does not become overconfident.
- "Bayesian or frequentist A/B testing?" — Bayesian gives a direct probability that B beats A and handles prior knowledge; frequentist gives p-values with well-known error guarantees. Both need a pre-agreed stopping rule.