Course Content
Statistics & Math for AI/ML Interviews
8 sections · 30 lessons
What is the difference between independent and dependent events in ML pipelines?
What you need to know
The intuition
Flip a coin twice. The first flip does not change the second, so the flips are independent. Now draw two cards from a deck without putting the first back. If the first card was an ace, there are fewer aces left, so the second draw is dependent on the first.
independent: P(A and B) = P(A) × P(B)dependent: P(A and B) = P(A) × P(B | A)P(B | A) is read "probability of B given A" — the next section covers it in depth.
Worked example. Chance of drawing two aces:
with replacement (independent): 4/52 × 4/52 = 1/169 ≈ 0.0059without replacement (dependent): 4/52 × 3/51 = 12/2652 ≈ 0.0045Where it matters in a pipeline
- Train/test splits. Random splitting assumes rows are independent. Rows from the same user, patient, device or document family are dependent, because they share hidden traits. The model can "recognise" the user instead of learning the pattern. This is a form of data leakage.
- Time series. Today's sales depend on yesterday's. A random split lets the model peek at the future. Split by time: train on the past, test on later data.
- Model assumptions. Naive Bayes assumes features are independent given the class. Many statistical tests assume independent observations; with repeated measures from the same user, their p-values are too optimistic.
The fix for grouped data is a group-aware split:
1import numpy as np2from sklearn.model_selection import GroupShuffleSplit34emails = np.arange(10) # 10 emails5campaign = np.array([1, 1, 1, 2, 2, 3, 3, 3, 4, 4]) # which campaign sent each one67splitter = GroupShuffleSplit(n_splits=1, test_size=0.3, random_state=0)8train_idx, test_idx = next(splitter.split(emails, groups=campaign))9print("train campaigns:", sorted(set(campaign[train_idx].tolist())))10print("test campaigns :", sorted(set(campaign[test_idx].tolist())))train campaigns: [1, 2]test campaigns : [3, 4]Every campaign lands entirely in train or entirely in test, so the test set measures performance on campaigns the model has never seen. GroupKFold does the same for cross-validation.
A real-life example
A team builds a spam filter. Spammers send the same message with tiny changes to thousands of addresses. With a random split, near-copies of each campaign sit in both train and test. The filter scores 98% on test. In production it catches far less, because each new campaign is genuinely new.
After switching to a split grouped by campaign, the offline score drops to 91% — and that number matches what production sees. The lower number is the honest one.
An everyday version: a teacher sets a test using questions the class already practised word for word. The scores are high, but they measure memory, not understanding. The questions and the practice were dependent.
Follow-up questions to expect
- "How do you test whether two events are independent?" — Check whether P(A and B) is close to P(A) × P(B), or whether P(B | A) is close to P(B). For categorical data, a chi-square test of independence does this formally.
- "What is leakage?" — Information in training that would not be available at prediction time, including overlap between train and test through dependent rows; it makes offline scores look better than reality.
- "Is uncorrelated the same as independent?" — No. Correlation only measures linear relationships; two variables can be uncorrelated and still strongly dependent, such as x and x squared for x centred on zero.