Statistics & Math for AI/ML Interviews

Course Content

Statistics & Math for AI/ML Interviews

8 sections · 30 lessons

Why is Bayes’ Theorem critical in probabilistic ML and NLP tasks?


What you need to know

Naive Bayes for text

To classify an email, Naive Bayes computes, for each class:

Text
P(class | words) is proportional to P(class) × P(word1 | class) × P(word2 | class) × ...

It learns P(word | class) by counting words in each class, which is fast and needs little data. "Naive" means it assumes words are independent given the class. Laplace smoothing (alpha=1) adds one to every count so an unseen word does not make a probability zero — itself a simple prior.

Python
from sklearn.feature_extraction.text import CountVectorizerfrom sklearn.naive_bayes import MultinomialNBtexts = ["win a free prize now", "free recharge offer claim now",         "claim your cashback prize", "meeting moved to monday",         "please review the invoice", "lunch at noon on monday"]labels = ["spam", "spam", "spam", "ham", "ham", "ham"]vec = CountVectorizer()model = MultinomialNB(alpha=1.0).fit(vec.fit_transform(texts), labels)for msg in ["claim your free prize", "invoice for monday meeting"]:    p_spam = model.predict_proba(vec.transform([msg]))[0][1]    print(f"{msg!r}: P(spam) = {p_spam:.3f}")
Text
'claim your free prize': P(spam) = 0.982'invoice for monday meeting': P(spam) = 0.077

Six training emails are enough to separate these two messages. Naive Bayes remains a strong, cheap baseline for text classification, even in the era of transformers.

Other places Bayes appears

  • Regularisation as a prior: L2 penalty = Gaussian prior on weights.
  • Bayesian optimisation: keeps a probabilistic model of "score as a function of hyperparameters", updates it after every trial, and picks the next trial where improvement is most likely.
  • Uncertainty: Bayesian neural networks and approximations such as MC dropout give a spread of predictions, not just one number.
  • Classic NLP: noisy-channel spelling correction picks the word w that maximises P(typed text | w) × P(w).

Correcting for a changed base rate

A fraud model is trained on a balanced dataset (50% fraud) but deployed where fraud is 1%. Its probabilities are too high. Bayes gives the fix: rescale the odds by the ratio of the new and old prior odds.

Python
def adjust_for_new_prior(p, train_prior, live_prior):    """Re-weight a probability when the base rate changes (prior shift)."""    odds = p / (1 - p)    odds *= (live_prior / (1 - live_prior)) / (train_prior / (1 - train_prior))    return odds / (1 + odds)print(round(adjust_for_new_prior(0.80, train_prior=0.5, live_prior=0.01), 3))
Text
0.039

A score of 0.80 on balanced training data means about 3.9% in production. Teams that skip this step set thresholds on inflated probabilities. (This correction assumes only the base rate changed, not the behaviour of fraud itself.)

A real-life example

An email provider's spam filter has 98% recall and a 1% false-positive rate. In one corporate customer's inbox, only 2% of mail is spam. Of 10,000 emails: 200 spam, 196 caught; 9,800 real, 98 wrongly flagged. So a third of the spam folder (98 of 294) is real mail. For a customer where 50% of mail is spam, only about 1% of the spam folder would be real mail. Same model, very different experience — because of the prior. The provider sets per-customer thresholds using each customer's measured spam rate.

Follow-up questions to expect

  • "Why is it called Naive Bayes?" — It assumes features are conditionally independent given the class, which is rarely true; it still ranks well because classification only needs the top class to win.
  • "Is a neural network classifier Bayesian?" — Its softmax output estimates a posterior P(class | input), but training finds one set of weights rather than a posterior over weights, so it is not fully Bayesian.
  • "What does Laplace smoothing do?" — Adds a small count to every word so unseen words do not force a probability of zero; it can be derived from a uniform prior over the word probabilities.