Statistics & Math for AI/ML Interviews

Course Content

Statistics & Math for AI/ML Interviews

8 sections · 30 lessons

What do prior, likelihood, and posterior represent in Bayesian models?


What you need to know

Text
posterior = likelihood × prior / evidenceP(H | D)  = P(D | H)   × P(H)  / P(D)

(H = hypothesis, D = data.)

TermSymbolQuestion it answersSpam example
PriorP(H)What did I believe before this data?20% of emails are spam
LikelihoodP(D given H)If H were true, how probable is this data?"cashback" appears in 10% of spam and 2% of ham
EvidenceP(D)How probable is this data overall?0.10 × 0.2 + 0.02 × 0.8 = 0.036
PosteriorP(H given D)What should I believe now?0.02 / 0.036 ≈ 0.56

Worked through: before reading the email you thought 20% spam. The word "cashback" is five times more common in spam than ham, so the belief rises to about 56%. Not 100% — one word is only one clue.

Where each term lives in ML

  • Prior: class frequencies in Naive Bayes; the prior over weights in Bayesian models. L2 regularisation is equivalent to putting a Gaussian prior on the weights and finding the most probable weights (MAP estimate); L1 corresponds to a Laplace prior.
  • Likelihood: the thing maximum-likelihood training maximises. Minimising cross-entropy loss is maximising the likelihood of the labels.
  • Posterior: the model's output P(class | input), or, in fully Bayesian models, a whole distribution over parameters that gives uncertainty estimates.

Strong versus weak priors

A prior can be weak (easily overturned by data) or strong. With little data, the prior dominates. With lots of data, the likelihood dominates and almost any reasonable prior gives the same answer.

A real-life example

A new batter scores 90 and 110 in their first two innings — an average of 100. Should the selectors believe their true average is 100? Almost nobody sustains that. Suppose past data says top-order batters typically average around 35.

A simple Bayesian-style estimate treats the prior as if it were worth, say, 10 innings of evidence at 35:

Text
estimate = (10 × 35 + 2 × 100) / (10 + 2) = 550 / 12 ≈ 46

The estimate is pulled toward the prior. After 50 innings, the batter's own record dominates. This shrinkage is how sensible rating systems handle new players, new restaurants with two five-star reviews, and new products with three sales. An e-commerce site that ranks by raw average rating puts a product with one 5-star review above one with 4.7 stars from 2,000 reviews; a prior fixes that.

Follow-up questions to expect

  • "What is the difference between MLE and MAP?" — Maximum likelihood picks parameters that make the data most probable; MAP also multiplies by a prior, which acts like regularisation.
  • "What is a conjugate prior?" — A prior that gives a posterior of the same family, making updates simple, such as a Beta prior for a conversion rate with Binomial data.
  • "What happens with a very strong wrong prior?" — It takes a lot of data to overturn it, so the posterior stays biased; that is why priors should be justified and checked.