Statistics & Math for AI/ML Interviews

Course Content

Statistics & Math for AI/ML Interviews

8 sections · 30 lessons

How does conditional probability relate to feature dependencies in ML models?


What you need to know

What dependence between features looks like

In an e-commerce dataset, P(customer is premium member | orders more than 10 per month) might be 0.6, while P(premium member) overall is 0.1. The two features are dependent: one carries information about the other.

Naive Bayes double-counts dependent features

Naive Bayes multiplies one likelihood per feature. That is only correct if the features are conditionally independent.

Worked example. Prior: 50% spam, so the prior odds are 1 to 1. The word "free" appears in 60% of spam and 20% of ham, a likelihood ratio of 0.6 / 0.2 = 3.

Text
one "free" feature:        odds = 1 × 3     = 3   →  P(spam) = 3/4  = 0.75"free" and "FREE!!!" astwo features that alwaysappear together:           odds = 1 × 3 × 3 = 9   →  P(spam) = 9/10 = 0.90

The second feature adds no new information, but Naive Bayes treats it as a second, independent clue. The ranking of emails often stays sensible, which is why Naive Bayes still works in practice, but its probabilities become overconfident.

Linear models: multicollinearity

When two features carry the same information, the model can trade weight between them freely, so coefficients swing between samples.

Python
import numpy as npfrom sklearn.linear_model import LinearRegressionrng = np.random.default_rng(0)n = 200area_sqft = rng.uniform(500, 2000, n)area_sqm = area_sqft / 10.764 + rng.normal(0, 2, n)   # same info, other unitrent = 20 * area_sqft + rng.normal(0, 3000, n)          # rupees per monthX = np.column_stack([area_sqft, area_sqm])for seed in [1, 2, 3]:    idx = np.random.default_rng(seed).choice(n, n)       # bootstrap sample    coef = LinearRegression().fit(X[idx], rent[idx]).coef_    print(f"sample {seed}: sqft coef = {coef[0]:7.1f}, sqm coef = {coef[1]:7.1f}")
Text
sample 1: sqft coef =    19.0, sqm coef =    15.9sample 2: sqft coef =    23.3, sqm coef =   -30.4sample 3: sqft coef =    18.0, sqm coef =    25.6

The square-metre coefficient flips from +25.6 to -30.4 across resamples. Yet the combined effect per square foot (sqft coef + sqm coef / 10.764) stays about 20.4 each time. Predictions are fine; the individual coefficients are meaningless. Ridge regularisation or dropping one of the two features fixes it.

Tree ensembles

Trees handle dependence well for prediction. But feature importance is split between the correlated features, so each can look unimportant even though together they matter a lot.

A real-life example

A spam team uses Naive Bayes with features for "free", "winner", "claim" and "prize". In lottery-scam emails these four words almost always appear together. Every such email gets a spam probability of 0.9999, including a genuine newsletter announcing a quiz winner. Because the probabilities are so extreme, no threshold separates "clearly spam" from "somewhat suspicious".

The team keeps Naive Bayes for its speed but calibrates the output scores on held-out data, and merges the four words into one "lottery language" feature. The probabilities become usable again.

Follow-up questions to expect

  • "How do you detect dependent features?" — Correlation matrices for numeric features, Cramér's V or mutual information for categorical ones, and the variance inflation factor (VIF) for multicollinearity in linear models.
  • "Why does Naive Bayes still work despite its assumption?" — Classification only needs the right class to have the highest score; double-counting distorts the probabilities more than the ranking.
  • "Is a feature that depends on the target always useful?" — For prediction yes, unless it is leakage — information only available after the outcome, such as "refund issued" in a fraud model.