Course Content
Machine Learning Foundations
14 sections · 70 lessons
What are common real-world ML mistakes freshers make in interviews and projects?
What you need to know
Interviewers ask this to see whether you have learned from real work. A list is weak; each mistake paired with its fix is strong.
| Mistake | Why it hurts | Fix |
|---|---|---|
| No baseline | "85%" means nothing without a reference | DummyClassifier, current business rule |
| Accuracy on imbalanced data | 99% while catching nothing | Precision, recall, PR-AUC, confusion matrix |
| Preprocessing before splitting | Test data leaks into training | Scaler and encoder inside a Pipeline |
| Features from the future | Great offline, useless live | Point-in-time features, a cutoff per row |
| Tuning on the test set | Test score becomes optimistic | Tune on validation or CV; test once |
| Random split on time or grouped data | Model memorises users or the future | Time-based split, GroupKFold |
| One aggregate metric only | Hides failing segments | Per-segment metrics, read the errors |
| No business framing | Wrong metric and threshold | Ask what each error costs |
| No reproducibility | Can't recreate the result | Fixed seeds, versioned data and config |
| No deployment thinking | Model breaks silently | Monitoring, retraining plan, rollback |
One mistake in detail: the grouped split
This one surprises many freshers because nothing looks wrong. Suppose each user has 20 transactions, and the label is per user. A random split puts some of each user's rows in training and the rest in testing, so the model can recognise the user rather than learn a pattern. Here the labels are random — there is nothing real to learn:
1import numpy as np2from sklearn.ensemble import RandomForestClassifier3from sklearn.model_selection import KFold, GroupKFold, cross_val_score45rng = np.random.default_rng(0)6users = 2007user_style = rng.normal(size=(users, 5)) # each user's own habits8user_label = rng.integers(0, 2, users) # label per user, unrelated to habits9user_id = np.repeat(np.arange(users), 20) # 20 transactions per user10X = user_style[user_id] + rng.normal(0, 0.1, (len(user_id), 5))11y = user_label[user_id]1213model = RandomForestClassifier(n_estimators=100, random_state=0, n_jobs=-1)14random_cv = cross_val_score(model, X, y, cv=KFold(5, shuffle=True, random_state=0))15group_cv = cross_val_score(model, X, y, cv=GroupKFold(5), groups=user_id)16print(f"random split accuracy: {random_cv.mean():.2f}")17print(f"grouped by user: {group_cv.mean():.2f}")random split accuracy: 0.99grouped by user: 0.5599% with a random split; 55% — close to a coin flip — when every user's rows stay on one side. The random split measured the model's memory of users it had already seen. For a model that will score new users, only the grouped number is honest. The same trap appears with patients who have several scans, and with photos of the same product.
Interview-specific mistakes
- Claiming more than you understand. If your résumé says "built a fraud model with XGBoost", expect "why XGBoost, which hyperparameters mattered, what was the precision at your threshold?".
- Listing tools instead of decisions. "I used pandas, scikit-learn and Flask" says little. "I chose recall at 20% precision because reviewers could handle 500 alerts a day" says a lot.
- Never mentioning what went wrong. Interviewers trust candidates who can describe a bug they found and fixed.
A real-life example
A final-year student presents a project: "Diabetes prediction with 98% accuracy." The interviewer asks three questions. What share of patients were diabetic? (8% — so "no one is diabetic" scores 92%.) Did you scale before splitting? (Yes.) Did any patient appear in both train and test? (The dataset had repeat visits, and the student had not checked.) After fixing all three, the honest result is recall of 71% at 40% precision. The student who can tell this story — the mistakes, the fixes and the honest number — is far more hireable than one who defends the 98%.
Follow-up questions to expect
- "Tell me about a mistake you made in an ML project." — Pick a real one, explain how you noticed it (a suspicious score, an error analysis), how you fixed it, and what habit you adopted afterwards.
- "How do you make an experiment reproducible?" — Fix random seeds, version the data and code, log hyperparameters and metrics for every run, and pin library versions.
- "How would you know your model is still working six months after launch?" — Monitor input distributions, prediction distribution and the business metric, and compare with labels as they arrive.