Course Content
Machine Learning Foundations
14 sections · 70 lessons
How do you debug an ML model that performs well offline but fails in production?
What you need to know
Offline, the model saw a carefully built dataset. Online, it sees whatever the live system sends it. When the two disagree, the model fails silently — it still returns a number, just the wrong one.
The usual causes, in the order to check them
- Training–serving skew — a feature is computed differently in production: km versus metres, missing values filled with 0 instead of the median, a daily aggregate that is 24 hours stale.
- Offline evaluation was too optimistic — leakage, or a random split on time-ordered data, so the offline score was never real.
- Data drift — the inputs changed: new users, a new city, a festive season (covariate shift).
- Concept drift — the relationship changed: fraudsters adapt, so the same features now mean something different.
- Operational issues — timeouts that trigger a fallback score, an old model version still deployed, a feature service returning nulls.
- Feedback loops — the model's own actions change the data it later sees; for example, blocked transactions never get a fraud label.
Checking for skew in code
The fastest check is to compare what the model saw in training with what it receives live. Log the live features (they are cheap to store) and compare simple statistics:
1import numpy as np, pandas as pd23rng = np.random.default_rng(0)4train = pd.DataFrame({"distance_km": rng.uniform(0.5, 8, 10_000),5 "prep_time_min": rng.normal(14, 4, 10_000).clip(3)})6live = pd.DataFrame({"distance_km": rng.uniform(0.5, 8, 2_000),7 "prep_time_min": rng.normal(14, 4, 2_000).clip(3)})8live.loc[:599, "distance_km"] *= 1000 # new app version sends metres9live.loc[1500:, "prep_time_min"] = 0 # feature service timeout, filled with 01011def profile(df):12 return pd.DataFrame({"median": df.median(), "p99": df.quantile(0.99),13 "share_zero": (df == 0).mean()})1415report = profile(train).join(profile(live), lsuffix="_train", rsuffix="_live")16print(report.round(2).T) distance_km prep_time_minmedian_train 4.26 14.00p99_train 7.93 23.20share_zero_train 0.00 0.00median_live 5.73 12.32p99_live 7702.25 22.70share_zero_live 0.00 0.25Two bugs jump out. The 99th percentile of distance went from 7.9 to 7,702 — some requests are in metres. And a quarter of live prep times are exactly 0, which never happens in training — a timeout default. Neither bug throws an error; both quietly ruin predictions. In production you would run this kind of comparison automatically and alert on it, often with a drift score such as the population stability index (PSI).
Prevention
- One feature code path for training and serving, or a feature store that serves the same values to both.
- Point-in-time correct training data, so offline features match what was known at prediction time.
- Time-based validation — train on the past, validate on a later period.
- Shadow deployment — run the new model on live traffic without acting on it, and compare.
- Monitoring — input feature distributions, prediction distribution, and the business metric, with alerts.
A real-life example
A food-delivery company launches a new ETA model that was 15% more accurate offline. In the first week, customer complaints about late orders rise. Logging shows the live model receives restaurant_load — the number of open orders at the restaurant — from a cache refreshed every 15 minutes, while training used the exact value at order time. During the dinner rush, the cached value is badly out of date, so the model thinks busy restaurants are quiet. The team moves the feature to a streaming source with a 30-second delay and adds the same delay when building training data. The offline score drops slightly, but now it matches production.
Follow-up questions to expect
- "What is the difference between data drift and concept drift?" — Data drift means the input distribution changes (more users from a new city). Concept drift means the relationship between inputs and the target changes (the same behaviour now signals fraud when it didn't before).
- "How do you monitor a model when labels arrive late?" — Monitor inputs and the prediction distribution immediately, use proxy metrics (chargebacks, complaints), and compute real metrics once labels arrive.
- "What is a shadow deployment?" — The new model receives real traffic and its predictions are logged, but the old model's predictions are used. You compare them before switching.