Course Content
Python for AI and Data Science
5 sections · 13 lessons
Model Training & Evaluation — Fitting, Scoring, and Tuning Models
A hospital asks for a screening model. The disease affects 1 in 100 people. You train a classifier, evaluate it, and report 99% accuracy. Everyone is impressed until someone checks what it predicts.
print(model.predict(X_test).sum()) # 0It has never predicted the disease. Not once. It learned that saying "healthy" every time is right 99% of the time, and by the metric you chose, it is excellent. It would miss every single patient who needed treatment.
This is not an exotic edge case. It is what happens by default whenever one outcome is rarer than the other, which is true of fraud, churn, defects, disease and nearly everything worth predicting. Training a model is the easy part — three lines. Knowing whether it works is the whole job.
Three kinds of problem, one interface
1from sklearn.ensemble import RandomForestClassifier, RandomForestRegressor2from sklearn.cluster import KMeans34clf = RandomForestClassifier(random_state=42).fit(X_train, y_train) # a category5reg = RandomForestRegressor(random_state=42).fit(X_train, y_train) # a number6km = KMeans(n_clusters=4, n_init=10, random_state=42).fit(X) # no labels| Problem | Predicts | Needs labels | Judged by |
|---|---|---|---|
| Classification | A class: spam, churn, digit | Yes | Precision, recall, F1, AUC |
| Regression | A number: price, temperature | Yes | MAE, RMSE, R² |
| Clustering | Group membership | No | Silhouette score; mostly human judgement |
The confusion matrix, and everything derived from it
Every classification metric is a different arithmetic summary of four counts. Here is a real model on 10,000 patients where 100 have the disease:
| Model says: has disease | Model says: healthy | Total | |
|---|---|---|---|
| Actually has disease | 80 (true positive) | 20 (false negative) | 100 |
| Actually healthy | 300 (false positive) | 9,600 (true negative) | 9,900 |
| Total | 380 | 9,620 | 10,000 |
Now compute the metrics by hand, because the arithmetic is what makes them memorable.
Accuracy — how often it was right overall:
Precision — of everyone it flagged, how many really had it:
Recall — of everyone who had it, how many it caught:
F1 — the harmonic mean, which punishes imbalance between the two:
Look at what those numbers say together. Accuracy of 96.8% sounds like a strong model. Precision of 21% says that four out of every five people it alarms are healthy. Recall of 80% says it catches four of every five real cases. Those are three completely different descriptions of the same model, and only one of them — the one you choose — will reach the person making the decision.
Accuracy answers "how often is it right?", which is almost never the question. The real question is "what does each kind of mistake cost?", and only precision and recall can answer it.
Choosing between precision and recall
You cannot maximise both. Lowering the decision threshold catches more true cases and also raises more false alarms; raising it does the reverse. Which way you should lean depends entirely on the relative cost of the two errors.
| Situation | Costly error | Optimise |
|---|---|---|
| Cancer screening | Missing a case (false negative) | Recall |
| Spam filter | Binning a real email (false positive) | Precision |
| Fraud blocking | Freezing a legitimate customer | Precision |
| Predictive maintenance | Missing a failure | Recall |
| Search results | An irrelevant top hit | Precision |
1from sklearn.metrics import classification_report, confusion_matrix23proba = model.predict_proba(X_test)[:, 1]4preds = (proba >= 0.30).astype(int) # move the threshold to favour recall56print(confusion_matrix(y_test, preds))7print(classification_report(y_test, preds, digits=3))The default threshold of 0.5 is a convention, not a law. It is very rarely the right operating point for imbalanced data, and choosing it deliberately is often a bigger improvement than any change of algorithm.
ROC-AUC and its imbalanced-data trap
ROC-AUC measures ranking quality across all thresholds: the probability that a randomly chosen positive is scored above a randomly chosen negative. 0.5 is random, 1.0 is perfect.
1from sklearn.metrics import roc_auc_score, average_precision_score23print(roc_auc_score(y_test, proba)) # 0.94 -- looks superb4print(average_precision_score(y_test, proba)) # 0.31 -- the honest numberThat gap is the trap. ROC-AUC's false positive rate has the enormous negative class in its denominator, so 300 false positives out of 9,900 healthy people barely moves it. Average precision — the area under the precision-recall curve — has only the positives in view and reports what you would actually experience.
| Metric | Use when | Misleads when |
|---|---|---|
| Accuracy | Classes are balanced and errors cost the same | Anything rare |
| Precision / recall / F1 | Almost always | You report only one of them |
| ROC-AUC | Comparing rankers on balanced data | The positive class is rare |
| Average precision | Rare positives — the default for imbalance | Rarely |
For multiclass problems, the averaging choice matters as much as the metric. average="macro" treats every class equally, so a rare class carries the same weight as a common one; average="weighted" weights by class size, letting a large class hide poor performance on small ones. If the rare classes are the point, use macro.
Regression metrics
1from sklearn.metrics import mean_absolute_error, mean_squared_error, r2_score2import numpy as np34y_true = np.array([10, 20, 30, 40, 1000])5y_pred = np.array([12, 18, 33, 38, 100])67print(mean_absolute_error(y_true, y_pred)) # 181.88print(np.sqrt(mean_squared_error(y_true, y_pred))) # 402.5Four of those five predictions are within 3 units. One is out by 900. MAE averages the errors and reports 181.8; RMSE squares them first, so the single 900 contributes 810,000 to a total of 810,021 and dominates entirely, giving 402.5.
That difference is the diagnostic. RMSE much larger than MAE means a small number of large errors; RMSE close to MAE means errors are evenly spread. Which you should report depends on whether one enormous error is worse than many small ones — for a delivery time estimate, probably yes; for a total revenue forecast, probably not.
| Metric | Units | Outlier sensitivity | Read as |
|---|---|---|---|
| MAE | Same as target | Low | Typical error size |
| RMSE | Same as target | High | Error size, with big misses punished |
| R² | None | High | Fraction of variance explained |
| MAPE | Percentage | Medium | Relative error — breaks if any true value is 0 |
R² compares your model against the simplest possible baseline: always predicting the mean. R² of 0.0 means you have matched that baseline and added nothing. R² can go negative, which means your model is worse than the mean — a result that reliably indicates a bug, usually a preprocessing step applied inconsistently between fitting and prediction.
Why one split is not enough
1from sklearn.model_selection import train_test_split2from sklearn.ensemble import RandomForestClassifier34for seed in range(5):5 Xtr, Xte, ytr, yte = train_test_split(X, y, test_size=0.2, random_state=seed)6 m = RandomForestClassifier(random_state=42).fit(Xtr, ytr)7 print(f"seed {seed}: {m.score(Xte, yte):.3f}")8# 0.847 0.881 0.859 0.902 0.864Same data, same model, same settings — and a spread of 5.5 percentage points depending on which rows happened to land in the test set. Report the 0.902 and you are reporting luck. This also means that if you compare two models on one split and they differ by 2 points, you have learned nothing.
Cross-validation
In the examples from here on, pipe is a pipeline like the one built in the previous lesson: preprocessing, then a model step named clf.
1from sklearn.model_selection import cross_val_score, StratifiedKFold23cv = StratifiedKFold(n_splits=5, shuffle=True, random_state=42)4scores = cross_val_score(pipe, X, y, cv=cv, scoring="f1_macro")5print(f"{scores.mean():.3f} ± {scores.std():.3f}") # 0.871 ± 0.021The data is divided into five parts; the model is trained five times, each time holding out a different part. Every row gets to be test data exactly once, and you get a mean plus a spread. The spread is not a footnote — it is what tells you whether a 0.006 difference between two models is real.
1from sklearn.model_selection import cross_validate23results = cross_validate(pipe, X, y, cv=cv,4 scoring=["precision_macro", "recall_macro", "f1_macro"],5 return_train_score=True)return_train_score=True gives you the overfitting gap for free: a training F1 of 0.99 against a validation F1 of 0.71 diagnoses the problem without any extra work.
Use StratifiedKFold for classification, plain KFold for regression, GroupKFold when rows cluster by customer or patient, and TimeSeriesSplit when the data is ordered in time — that last one trains only on earlier folds and tests on later ones, because shuffling a timeline means training on the future.
Tuning hyperparameters
Hyperparameters are the settings you choose rather than the ones the model learns: tree depth, number of neighbours, regularisation strength. Search them systematically rather than by hand.
1from sklearn.model_selection import GridSearchCV, RandomizedSearchCV2from scipy.stats import randint, uniform34grid = {5 "clf__n_estimators": [100, 300, 500],6 "clf__max_depth": [None, 10, 20],7 "clf__min_samples_leaf": [1, 5, 10],8}9search = GridSearchCV(pipe, grid, cv=cv, scoring="f1_macro", n_jobs=-1, refit=True)10search.fit(X_train, y_train)1112print(search.best_params_)13print(f"{search.best_score_:.3f}")14final = search.best_estimator_ # already refitted on all training dataNote the double underscore in clf__max_depth. That is how you address a parameter of a named step inside a pipeline — step name, two underscores, parameter name — and it is what allows the scaler to be refitted inside every fold of every candidate configuration.
| Search | Fits required | Best for |
|---|---|---|
GridSearchCV | Every combination × folds | Few parameters with few sensible values |
RandomizedSearchCV | n_iter × folds — you choose | Many parameters, or continuous ranges |
HalvingRandomSearchCV | Far fewer | Large budgets; eliminates weak candidates early |
The grid above is 3 × 3 × 3 = 27 combinations × 5 folds = 135 model fits. Add one more parameter with four values and it is 540. Grid search grows multiplicatively, so beyond three parameters random search finds better configurations for the same compute — most parameters barely matter, and random sampling spends its budget exploring the ones that do rather than exhaustively mapping the ones that do not.
1dist = {2 "clf__n_estimators": randint(100, 800),3 "clf__max_depth": randint(3, 30),4 "clf__min_samples_leaf": randint(1, 20),5 "clf__max_features": uniform(0.1, 0.9),6}7search = RandomizedSearchCV(pipe, dist, n_iter=60, cv=cv,8 scoring="f1_macro", n_jobs=-1, random_state=42)The score reported by a hyperparameter search is optimistic — you selected the best of many attempts. Only an untouched test set gives an honest final number.
Diagnosing what is wrong
| Training score | Validation score | Diagnosis | Fix |
|---|---|---|---|
| High (0.99) | Much lower (0.71) | Overfitting | More data, simpler model, regularisation, fewer features |
| Low (0.65) | Low (0.63) | Underfitting | More capacity, better features, less regularisation |
| High | High | Working — or leaking | Check for leakage before celebrating |
| Low | Higher than training | Something is wrong | Usually a broken split or an inconsistent transform |
1from sklearn.model_selection import learning_curve2import matplotlib.pyplot as plt3import numpy as np45sizes, train_sc, val_sc = learning_curve(6 pipe, X, y, cv=cv, scoring="f1_macro",7 train_sizes=np.linspace(0.1, 1.0, 8), n_jobs=-1)89plt.plot(sizes, train_sc.mean(1), "o-", label="training")10plt.plot(sizes, val_sc.mean(1), "o-", label="validation")11plt.xlabel("Training examples"); plt.ylabel("F1"); plt.legend()The learning curve answers a question worth real money: will more data help? If the two lines are converging and the gap is still closing, collecting more data will improve the model. If they have flattened and met at a low value, more data of the same kind will change nothing, and your effort belongs in features or in a different model family.
Feature importance, and the bias in the obvious version
1import pandas as pd2from sklearn.inspection import permutation_importance34# one value per column AFTER preprocessing (each one-hot level separately)5built_in = pd.Series(final["clf"].feature_importances_,6 index=final[:-1].get_feature_names_out())78# one value per ORIGINAL column: the whole pipeline is scored9perm = permutation_importance(final, X_test, y_test, n_repeats=20,10 random_state=42, scoring="f1_macro")11honest = pd.Series(perm.importances_mean, index=X_test.columns)1213print(built_in.sort_values(ascending=False).head(10).round(4))14print(honest.sort_values(ascending=False).round(4))A tree's built-in importance counts how much each feature reduced impurity during training, which systematically favours features with many distinct values — a customer ID or a continuous variable offers more places to split, so it accumulates importance whether or not it generalises. Permutation importance instead shuffles one column in the test set and measures how much the score drops. If shuffling a feature does not hurt performance, the model was not really using it. The two lists are at different levels when the pipeline one-hot encodes: the built-in version scores each encoded column (cat__region_north), while permutation scores the original column (region) as a whole.
Where the two disagree sharply, believe the permutation version. A feature that ranks first by impurity and near zero by permutation is a feature the model memorised rather than learned from.
What to report when you are done
1from sklearn.metrics import classification_report, confusion_matrix, f1_score23final = search.best_estimator_4preds = final.predict(X_test) # the test set, touched once, at the end56print(confusion_matrix(y_test, preds))7print(classification_report(y_test, preds, digits=3))8print(f"cross-validated: {search.best_score_:.3f}")9print(f"held-out test: {f1_score(y_test, preds, average='macro'):.3f}")Report four things, always. The metric and why you chose it, tied to what the two error types cost. The cross-validated mean and standard deviation, so the reader knows how much of any difference is noise. The held-out test score, computed once. And the confusion matrix, because it is the only output that shows the errors themselves rather than an average of them.
A model is not finished when it fits. It is finished when you can say what it gets wrong, how often, and what that costs — and when the number you are quoting came from data the model has never influenced in any way.
Check your understanding
0 of 3 answered
1.A fraud model reports 99.5% accuracy on data where 0.5% of transactions are fraud. What should you look at before trusting it?
2.A regression model has an MAE of 5 and an RMSE of 40. What does that gap tell you?
3.GridSearchCV reports a best cross-validated F1 of 0.91. Which number should go in your final report as the estimate of real performance?