Machine Learning Foundations

Course Content

Machine Learning Foundations

14 sections · 70 lessons

What are common real-world ML mistakes freshers make in interviews and projects?


200 users, 20 transactions each, random labelsRandom split• A user's rows land on both sides• Model recognises the user• Accuracy 0.99• Measures memory, not skillGroupKFold by user• Each user on one side only• Model meets only new users• Accuracy 0.55• Honest: nothing to learn
When rows share an owner, a random split lets the model recognise the owner instead of learning a pattern.

What you need to know

Interviewers ask this to see whether you have learned from real work. A list is weak; each mistake paired with its fix is strong.

MistakeWhy it hurtsFix
No baseline"85%" means nothing without a referenceDummyClassifier, current business rule
Accuracy on imbalanced data99% while catching nothingPrecision, recall, PR-AUC, confusion matrix
Preprocessing before splittingTest data leaks into trainingScaler and encoder inside a Pipeline
Features from the futureGreat offline, useless livePoint-in-time features, a cutoff per row
Tuning on the test setTest score becomes optimisticTune on validation or CV; test once
Random split on time or grouped dataModel memorises users or the futureTime-based split, GroupKFold
One aggregate metric onlyHides failing segmentsPer-segment metrics, read the errors
No business framingWrong metric and thresholdAsk what each error costs
No reproducibilityCan't recreate the resultFixed seeds, versioned data and config
No deployment thinkingModel breaks silentlyMonitoring, retraining plan, rollback

One mistake in detail: the grouped split

This one surprises many freshers because nothing looks wrong. Suppose each user has 20 transactions, and the label is per user. A random split puts some of each user's rows in training and the rest in testing, so the model can recognise the user rather than learn a pattern. Here the labels are random — there is nothing real to learn:

Python
import numpy as npfrom sklearn.ensemble import RandomForestClassifierfrom sklearn.model_selection import KFold, GroupKFold, cross_val_scorerng = np.random.default_rng(0)users = 200user_style = rng.normal(size=(users, 5))           # each user's own habitsuser_label = rng.integers(0, 2, users)             # label per user, unrelated to habitsuser_id = np.repeat(np.arange(users), 20)          # 20 transactions per userX = user_style[user_id] + rng.normal(0, 0.1, (len(user_id), 5))y = user_label[user_id]model = RandomForestClassifier(n_estimators=100, random_state=0, n_jobs=-1)random_cv = cross_val_score(model, X, y, cv=KFold(5, shuffle=True, random_state=0))group_cv = cross_val_score(model, X, y, cv=GroupKFold(5), groups=user_id)print(f"random split accuracy:  {random_cv.mean():.2f}")print(f"grouped by user:        {group_cv.mean():.2f}")
Text
random split accuracy:  0.99grouped by user:        0.55

99% with a random split; 55% — close to a coin flip — when every user's rows stay on one side. The random split measured the model's memory of users it had already seen. For a model that will score new users, only the grouped number is honest. The same trap appears with patients who have several scans, and with photos of the same product.

Interview-specific mistakes

  • Claiming more than you understand. If your résumé says "built a fraud model with XGBoost", expect "why XGBoost, which hyperparameters mattered, what was the precision at your threshold?".
  • Listing tools instead of decisions. "I used pandas, scikit-learn and Flask" says little. "I chose recall at 20% precision because reviewers could handle 500 alerts a day" says a lot.
  • Never mentioning what went wrong. Interviewers trust candidates who can describe a bug they found and fixed.

A real-life example

A final-year student presents a project: "Diabetes prediction with 98% accuracy." The interviewer asks three questions. What share of patients were diabetic? (8% — so "no one is diabetic" scores 92%.) Did you scale before splitting? (Yes.) Did any patient appear in both train and test? (The dataset had repeat visits, and the student had not checked.) After fixing all three, the honest result is recall of 71% at 40% precision. The student who can tell this story — the mistakes, the fixes and the honest number — is far more hireable than one who defends the 98%.

Follow-up questions to expect

  • "Tell me about a mistake you made in an ML project." — Pick a real one, explain how you noticed it (a suspicious score, an error analysis), how you fixed it, and what habit you adopted afterwards.
  • "How do you make an experiment reproducible?" — Fix random seeds, version the data and code, log hyperparameters and metrics for every run, and pin library versions.
  • "How would you know your model is still working six months after launch?" — Monitor input distributions, prediction distribution and the business metric, and compare with labels as they arrive.