Machine Learning Foundations

Course Content

Machine Learning Foundations

14 sections · 70 lessons

How do you debug a Machine Learning model with poor performance?


What you need to know

"The model is bad" is a symptom, not a diagnosis. An interviewer is testing whether you have a process — an order of checks that finds the cause quickly — rather than a habit of trying random algorithms.

  1. Check the framing and metric — is the target defined correctly? Is the metric suitable for the class balance and the business cost?
  2. Look at the data — missing values, duplicates, impossible values, label errors, and train versus production differences.
  3. Build a baseline — the simplest possible predictor. If your model barely beats it, something basic is wrong.
  4. Check for leakage — suspiciously good scores are a bug too.
  5. Diagnose bias versus variance — compare training and validation scores.
  6. Read the errors — look at the worst predictions and the confusion matrix; patterns point to missing features.
  7. Then tune and try other models — once the data and framing are sound.

Steps 2 and 3 in code

Here is a delivery-time dataset with two realistic problems: 5% of distances are missing, and a logging bug recorded 40 deliveries as taking 0 minutes.

Python
import numpy as np, pandas as pdfrom sklearn.dummy import DummyRegressorfrom sklearn.ensemble import HistGradientBoostingRegressorfrom sklearn.model_selection import cross_val_scorerng = np.random.default_rng(0)n = 3000df = pd.DataFrame({"distance_km": rng.uniform(0.5, 8, n),                   "is_peak": rng.integers(0, 2, n)})df["minutes"] = 15 + 3 * df.distance_km + 12 * df.is_peak + rng.normal(0, 3, n)df.loc[rng.choice(n, 150, replace=False), "distance_km"] = np.nan   # 5% missingdf.loc[rng.choice(n, 40, replace=False), "minutes"] = 0              # logging bug: 0-minute deliveriesprint("missing share:", df.isna().mean().round(3).to_dict())print("rows with minutes <= 0:", (df.minutes <= 0).sum())for label, data in [("raw", df), ("cleaned", df[df.minutes > 0])]:    X, y = data[["distance_km", "is_peak"]], data.minutes    for name, m in [("baseline", DummyRegressor()), ("boosting", HistGradientBoostingRegressor())]:        mae = -cross_val_score(m, X, y, cv=5, scoring="neg_mean_absolute_error").mean()        print(f"{label:7s} {name:8s} MAE = {mae:.1f} min")
Text
missing share: {'distance_km': 0.05, 'is_peak': 0.0, 'minutes': 0.0}rows with minutes <= 0: 40raw     baseline MAE = 7.8 minraw     boosting MAE = 3.3 mincleaned baseline MAE = 7.5 mincleaned boosting MAE = 2.7 min

Two quick checks found the problems. The baseline (DummyRegressor predicts the average) shows what "no skill" looks like: 7.8 minutes. The model clearly beats it, so the model is learning something. Removing 40 impossible labels — just 1.3% of rows — cuts the model's error from 3.3 to 2.7 minutes, because those zeros were pulling predictions down and inflating the error. No hyperparameter change could have found that. (HistGradientBoostingRegressor handles the missing distances natively; with other models you would impute them inside a pipeline.)

Step 5: the bias–variance read

Training scoreValidation scoreLikely causeDirection
LowLow, closeUnderfittingMore features, more flexible model, less regularisation
HighMuch lowerOverfittingMore data, regularisation, simpler model
Very highVery high, on a hard problemLeakageAudit features and the split

Step 6: read the errors

Sort validation rows by error and read the worst 50. In a delivery model you might find that almost all large errors are orders from one new restaurant, or orders placed during rain. That tells you which feature to add — far more useful than one more percentage point from tuning.

A real-life example

A grocery app's model predicts whether a customer will accept a substitute when an item is out of stock. It scores 61% accuracy, and the team is about to try a neural network. A senior engineer follows the checklist instead. The baseline "always accept" scores 58% — so the model is barely learning. Reading the data, she finds the label is recorded when the picker marks the substitute, not when the customer responds, so half the "accepted" labels are just default clicks. After fixing the label definition with the product team, the same simple model reaches 78%.

Follow-up questions to expect

  • "What baseline would you use?" — The majority class or the mean for a start, then the current business rule or heuristic, then a simple model like logistic regression.
  • "What if the model cannot beat the baseline?" — Suspect the data: wrong labels, a join that misaligned rows, features that were all null in training, or a target with no learnable signal.
  • "How do you find label errors?" — Look at confident wrong predictions: if the model is 99% sure and the label disagrees, the label is often wrong. Tools like cleanlab automate this idea.