Course Content
Statistics & Math for AI/ML Interviews
8 sections · 30 lessons
How does conditional probability relate to feature dependencies in ML models?
What you need to know
What dependence between features looks like
In an e-commerce dataset, P(customer is premium member | orders more than 10 per month) might be 0.6, while P(premium member) overall is 0.1. The two features are dependent: one carries information about the other.
Naive Bayes double-counts dependent features
Naive Bayes multiplies one likelihood per feature. That is only correct if the features are conditionally independent.
Worked example. Prior: 50% spam, so the prior odds are 1 to 1. The word "free" appears in 60% of spam and 20% of ham, a likelihood ratio of 0.6 / 0.2 = 3.
one "free" feature: odds = 1 × 3 = 3 → P(spam) = 3/4 = 0.75"free" and "FREE!!!" astwo features that alwaysappear together: odds = 1 × 3 × 3 = 9 → P(spam) = 9/10 = 0.90The second feature adds no new information, but Naive Bayes treats it as a second, independent clue. The ranking of emails often stays sensible, which is why Naive Bayes still works in practice, but its probabilities become overconfident.
Linear models: multicollinearity
When two features carry the same information, the model can trade weight between them freely, so coefficients swing between samples.
1import numpy as np2from sklearn.linear_model import LinearRegression34rng = np.random.default_rng(0)5n = 2006area_sqft = rng.uniform(500, 2000, n)7area_sqm = area_sqft / 10.764 + rng.normal(0, 2, n) # same info, other unit8rent = 20 * area_sqft + rng.normal(0, 3000, n) # rupees per month910X = np.column_stack([area_sqft, area_sqm])11for seed in [1, 2, 3]:12 idx = np.random.default_rng(seed).choice(n, n) # bootstrap sample13 coef = LinearRegression().fit(X[idx], rent[idx]).coef_14 print(f"sample {seed}: sqft coef = {coef[0]:7.1f}, sqm coef = {coef[1]:7.1f}")sample 1: sqft coef = 19.0, sqm coef = 15.9sample 2: sqft coef = 23.3, sqm coef = -30.4sample 3: sqft coef = 18.0, sqm coef = 25.6The square-metre coefficient flips from +25.6 to -30.4 across resamples. Yet the combined effect per square foot (sqft coef + sqm coef / 10.764) stays about 20.4 each time. Predictions are fine; the individual coefficients are meaningless. Ridge regularisation or dropping one of the two features fixes it.
Tree ensembles
Trees handle dependence well for prediction. But feature importance is split between the correlated features, so each can look unimportant even though together they matter a lot.
A real-life example
A spam team uses Naive Bayes with features for "free", "winner", "claim" and "prize". In lottery-scam emails these four words almost always appear together. Every such email gets a spam probability of 0.9999, including a genuine newsletter announcing a quiz winner. Because the probabilities are so extreme, no threshold separates "clearly spam" from "somewhat suspicious".
The team keeps Naive Bayes for its speed but calibrates the output scores on held-out data, and merges the four words into one "lottery language" feature. The probabilities become usable again.
Follow-up questions to expect
- "How do you detect dependent features?" — Correlation matrices for numeric features, Cramér's V or mutual information for categorical ones, and the variance inflation factor (VIF) for multicollinearity in linear models.
- "Why does Naive Bayes still work despite its assumption?" — Classification only needs the right class to have the highest score; double-counting distorts the probabilities more than the ranking.
- "Is a feature that depends on the target always useful?" — For prediction yes, unless it is leakage — information only available after the outcome, such as "refund issued" in a fraud model.