Course Content
Machine Learning Foundations
14 sections · 70 lessons
Why is model evaluation critical in Machine Learning?
What you need to know
A model is trained to do well on its training data. The question you actually care about is different: how well will it do on data it has never seen — next week's orders, tomorrow's payments? That ability is called generalisation, and evaluation is how you measure it.
Training score can be made perfect
Here a decision tree is trained twice on data where 10% of labels are noisy. One tree is limited to depth 4; the other can grow as deep as it wants.
1from sklearn.datasets import make_classification2from sklearn.model_selection import train_test_split3from sklearn.tree import DecisionTreeClassifier45X, y = make_classification(n_samples=2000, n_features=20, n_informative=5,6 flip_y=0.1, random_state=0) # 10% of labels are noisy7X_tr, X_te, y_tr, y_te = train_test_split(X, y, test_size=0.25, random_state=0)89for depth in [4, None]:10 tree = DecisionTreeClassifier(max_depth=depth, random_state=0).fit(X_tr, y_tr)11 print(f"max_depth={str(depth):4s} train={tree.score(X_tr, y_tr):.2f} test={tree.score(X_te, y_te):.2f}")max_depth=4 train=0.87 test=0.84max_depth=None train=1.00 test=0.82The unlimited tree scores a perfect 1.00 on training data. It has memorised every row, including the 10% of wrong labels. On the test set it is actually worse than the small tree. If you had looked only at the training score, you would have shipped the wrong model.
What evaluation lets you do
- Estimate real-world performance — the test score is your best guess of production behaviour.
- Compare models fairly — every candidate is scored on the same held-out data with the same metric.
- Choose settings — hyperparameters and the decision threshold are picked using validation data.
- Decide go or no-go — is the model better than the current system, or a simple rule?
- Monitor after launch — the same metric, tracked over time, tells you when the model degrades.
Good evaluation habits
- Pick the metric before training, so you are not tempted to choose whichever looks best afterwards.
- Make the test data look like production. If the model will predict next month, test on a later time period, not a random shuffle.
- Check segments, not just the average. A model can be 90% accurate overall and 60% accurate for new users.
- Compare with a baseline — "always predict the majority class" or the current business rule.
- Keep the test set locked until the final check (the Data Splitting section explains why).
A real-life example
A lending startup builds a model to predict which personal-loan applicants will default. The data scientist reports 94% accuracy. The credit head asks two questions: "94% compared with what?" and "which 6% is wrong?"
It turns out only 5% of applicants default, so "approve everyone" is already 95% accurate. The confusion matrix shows the model misses 70% of defaulters. The team switches to recall of defaulters and expected loss in rupees as the main metrics, and compares every model against the current rule-based scorecard. The "94% model" is never shipped.
Follow-up questions to expect
- "What is the difference between a validation set and a test set?" — The validation set is used repeatedly to choose models and settings. The test set is used once, at the end, for an unbiased final estimate.
- "Is a good test score enough to deploy?" — No. You also check performance on important segments, latency, fairness, and compare with the current system, often with a shadow or A/B test.
- "What is offline versus online evaluation?" — Offline is scoring on stored data; online is measuring business outcomes with real users, such as an A/B test on conversion rate.