Course Content
Statistics & Math for AI/ML Interviews
8 sections · 30 lessons
Can two features be highly correlated but have no causal relationship? Explain with an ML example.
What you need to know
Symptoms versus causes
Many strong ML features are symptoms of an underlying state rather than causes of the outcome.
login problems / frustration → more password resetslogin problems / frustration → churnThe two arrows share a source. Password resets and churn are correlated, but neither causes the other. Change one, and the other does not move.
Why models learn such features
A model minimises its loss. It uses whatever signal helps, with no notion of cause. This is fine until:
- Someone intervenes on the feature. The link breaks, because the feature was never the lever.
- The environment changes. If a new login flow reduces password resets for everyone, the model's churn predictions become wrong even though frustration has not changed.
- The feature is a shortcut. Image models have been shown to use hospital-specific markers on X-rays instead of the medical signs, which works in one hospital and fails in another.
A well-known healthcare example: a model predicting pneumonia death risk learned that patients with asthma had lower risk. The reason was that asthma patients were sent straight to intensive care and got more aggressive treatment. The correlation was real in the data, but using the model to send asthma patients home would have been dangerous.
How to handle these features
- Keep them for prediction if they are available at prediction time.
- Do not use their importance or coefficients to recommend interventions.
- Test interventions with an experiment.
- Monitor them: symptom features are often the first to drift when processes change.
A real-life example
A telecom churn model's top feature is "called customer care in the last 30 days". The operations head proposes making customer care harder to reach, "since calls lead to churn". The data scientist explains that dissatisfied customers call care, and dissatisfaction causes churn.
They instead test the opposite in an experiment: for half of the high-risk callers, care agents get authority to offer a small credit. Churn in that group falls from 18% to 14%. The model's correlation found the right customers; only the experiment found the right action.
A cricket version: a team notices it wins more often when its opening bowler takes an early wicket. Picking a bowler "who takes early wickets" sounds causal, but early wickets are also more likely on a green, seaming pitch, which helps the whole bowling attack. The pitch is part of the story.
Follow-up questions to expect
- "Should you remove non-causal features from a predictive model?" — Not necessarily. Keep them if they are available at prediction time, stable and not unfair proxies; remove them if they leak future information or encode bias.
- "How would you find a causal driver of churn?" — Form a hypothesis, then run a randomised experiment on the lever you control, such as offers, onboarding changes or pricing.
- "What is a spurious correlation?" — A correlation with no direct causal link, produced by a confounder, selection bias or chance.