Course Content
Machine Learning Foundations
14 sections · 70 lessons
How do you debug a Machine Learning model with poor performance?
What you need to know
"The model is bad" is a symptom, not a diagnosis. An interviewer is testing whether you have a process — an order of checks that finds the cause quickly — rather than a habit of trying random algorithms.
- Check the framing and metric — is the target defined correctly? Is the metric suitable for the class balance and the business cost?
- Look at the data — missing values, duplicates, impossible values, label errors, and train versus production differences.
- Build a baseline — the simplest possible predictor. If your model barely beats it, something basic is wrong.
- Check for leakage — suspiciously good scores are a bug too.
- Diagnose bias versus variance — compare training and validation scores.
- Read the errors — look at the worst predictions and the confusion matrix; patterns point to missing features.
- Then tune and try other models — once the data and framing are sound.
Steps 2 and 3 in code
Here is a delivery-time dataset with two realistic problems: 5% of distances are missing, and a logging bug recorded 40 deliveries as taking 0 minutes.
1import numpy as np, pandas as pd2from sklearn.dummy import DummyRegressor3from sklearn.ensemble import HistGradientBoostingRegressor4from sklearn.model_selection import cross_val_score56rng = np.random.default_rng(0)7n = 30008df = pd.DataFrame({"distance_km": rng.uniform(0.5, 8, n),9 "is_peak": rng.integers(0, 2, n)})10df["minutes"] = 15 + 3 * df.distance_km + 12 * df.is_peak + rng.normal(0, 3, n)11df.loc[rng.choice(n, 150, replace=False), "distance_km"] = np.nan # 5% missing12df.loc[rng.choice(n, 40, replace=False), "minutes"] = 0 # logging bug: 0-minute deliveries1314print("missing share:", df.isna().mean().round(3).to_dict())15print("rows with minutes <= 0:", (df.minutes <= 0).sum())1617for label, data in [("raw", df), ("cleaned", df[df.minutes > 0])]:18 X, y = data[["distance_km", "is_peak"]], data.minutes19 for name, m in [("baseline", DummyRegressor()), ("boosting", HistGradientBoostingRegressor())]:20 mae = -cross_val_score(m, X, y, cv=5, scoring="neg_mean_absolute_error").mean()21 print(f"{label:7s} {name:8s} MAE = {mae:.1f} min")missing share: {'distance_km': 0.05, 'is_peak': 0.0, 'minutes': 0.0}rows with minutes <= 0: 40raw baseline MAE = 7.8 minraw boosting MAE = 3.3 mincleaned baseline MAE = 7.5 mincleaned boosting MAE = 2.7 minTwo quick checks found the problems. The baseline (DummyRegressor predicts the average) shows what "no skill" looks like: 7.8 minutes. The model clearly beats it, so the model is learning something. Removing 40 impossible labels — just 1.3% of rows — cuts the model's error from 3.3 to 2.7 minutes, because those zeros were pulling predictions down and inflating the error. No hyperparameter change could have found that. (HistGradientBoostingRegressor handles the missing distances natively; with other models you would impute them inside a pipeline.)
Step 5: the bias–variance read
| Training score | Validation score | Likely cause | Direction |
|---|---|---|---|
| Low | Low, close | Underfitting | More features, more flexible model, less regularisation |
| High | Much lower | Overfitting | More data, regularisation, simpler model |
| Very high | Very high, on a hard problem | Leakage | Audit features and the split |
Step 6: read the errors
Sort validation rows by error and read the worst 50. In a delivery model you might find that almost all large errors are orders from one new restaurant, or orders placed during rain. That tells you which feature to add — far more useful than one more percentage point from tuning.
A real-life example
A grocery app's model predicts whether a customer will accept a substitute when an item is out of stock. It scores 61% accuracy, and the team is about to try a neural network. A senior engineer follows the checklist instead. The baseline "always accept" scores 58% — so the model is barely learning. Reading the data, she finds the label is recorded when the picker marks the substitute, not when the customer responds, so half the "accepted" labels are just default clicks. After fixing the label definition with the product team, the same simple model reaches 78%.
Follow-up questions to expect
- "What baseline would you use?" — The majority class or the mean for a start, then the current business rule or heuristic, then a simple model like logistic regression.
- "What if the model cannot beat the baseline?" — Suspect the data: wrong labels, a join that misaligned rows, features that were all null in training, or a target with no learnable signal.
- "How do you find label errors?" — Look at confident wrong predictions: if the model is 99% sure and the label disagrees, the label is often wrong. Tools like cleanlab automate this idea.