Course Content
Machine Learning Foundations
14 sections · 70 lessons
What is training data vs inference data?
What you need to know
Two phases of a model's life
- Training happens offline, usually once in a while. You have past records with outcomes, such as loans from 2022 to 2024 with "repaid" or "defaulted". The model learns from them.
- Inference happens every time the model is used. A new applicant fills a form today, and the model predicts default risk. Nobody knows the true outcome yet; it arrives months later, if at all.
The labels arrive late
At inference time you usually have no label. For loan default the label appears after 6 to 12 months. For churn, after the customer leaves or stays. This delay matters: it means you cannot immediately tell if the model is wrong in production, so you must monitor the inputs as well as the outputs.
Training–serving skew
The model only knows the world as it looked in training. If inference data is prepared differently, the model receives numbers it has never seen and gives confident nonsense. Here the house-price model from the previous lesson is trained on area in square feet, and a new app screen sends square metres:
1import joblib2model = joblib.load("house_model.joblib")34# The training data measured area in square feet.5print(model.predict([[1200, 2]]).round(1)) # correct units6# A new app screen sends area in square metres (1200 sq ft = 111.5 sq m)7print(model.predict([[111.5, 2]]).round(1)) # same flat, wrong units[85.5][16.3]Same flat, same model, and the price drops from 85.5 lakh to 16.3 lakh. Nothing crashes, no error is raised. That is why skew is dangerous: it fails silently.
Common causes of skew:
- Different code for features in training (a SQL query) and in serving (a Python function).
- A scaler or encoder re-fitted at serving time instead of reusing the one saved from training.
- Different defaults for missing values, such as 0 in training and -1 in serving.
- Different time windows, such as "spend in the last 30 days" versus "last 28 days".
Distribution shift
Even with identical code, the world changes. A model trained on pre-festival shopping data sees very different baskets during Diwali. When inference data drifts away from training data, accuracy drops and the model needs retraining.
A real-life example
A telecom company trained a churn model with the feature "number of complaint calls in the last 30 days", computed in a data warehouse from the full call log. In production, the app computed the same feature from a cache that only kept 7 days of calls. Every customer looked happier than they were, and the model predicted almost no churn. Offline recall was 0.78; live recall fell to about 0.30. The fix was to compute the feature once, in one shared function (a feature store serves this purpose), used by both training and serving.
Follow-up questions to expect
- "How do you prevent training–serving skew?" — Use one shared feature pipeline for both, save the fitted preprocessing with the model (a scikit-learn
Pipeline), and log live inputs to compare with training statistics. - "What is data drift?" — When the distribution of live inputs moves away from the training data, for example average order value rising after a price change. It is detected by comparing feature statistics over time.
- "Is test data the same as inference data?" — No. Test data is held-out labelled data used to estimate performance before launch. Inference data is live and usually unlabelled.