Python for AI and Data Science

Model Training & Evaluation — Fitting, Scoring, and Tuning Models


A hospital asks for a screening model. The disease affects 1 in 100 people. You train a classifier, evaluate it, and report 99% accuracy. Everyone is impressed until someone checks what it predicts.

Python
print(model.predict(X_test).sum())     # 0

It has never predicted the disease. Not once. It learned that saying "healthy" every time is right 99% of the time, and by the metric you chose, it is excellent. It would miss every single patient who needed treatment.

This is not an exotic edge case. It is what happens by default whenever one outcome is rarer than the other, which is true of fraud, churn, defects, disease and nearly everything worth predicting. Training a model is the easy part — three lines. Knowing whether it works is the whole job.

The 99% screening model, on 1,000 patients0100990Predicted illPredicted wellActually illActually wellPredicting well for everyone scores 99% accuracy and finds nobody.
Accuracy rewards the majority class; recall on the ten cases in the top row is the number the hospital asked for.

Three kinds of problem, one interface

Python
from sklearn.ensemble import RandomForestClassifier, RandomForestRegressorfrom sklearn.cluster import KMeansclf = RandomForestClassifier(random_state=42).fit(X_train, y_train)   # a categoryreg = RandomForestRegressor(random_state=42).fit(X_train, y_train)    # a numberkm  = KMeans(n_clusters=4, n_init=10, random_state=42).fit(X)         # no labels
ProblemPredictsNeeds labelsJudged by
ClassificationA class: spam, churn, digitYesPrecision, recall, F1, AUC
RegressionA number: price, temperatureYesMAE, RMSE, R²
ClusteringGroup membershipNoSilhouette score; mostly human judgement

The confusion matrix, and everything derived from it

Every classification metric is a different arithmetic summary of four counts. Here is a real model on 10,000 patients where 100 have the disease:

Model says: has diseaseModel says: healthyTotal
Actually has disease80 (true positive)20 (false negative)100
Actually healthy300 (false positive)9,600 (true negative)9,900
Total3809,62010,000

Now compute the metrics by hand, because the arithmetic is what makes them memorable.

Accuracy — how often it was right overall:

accuracy=80+960010000=0.968\text{accuracy} = \frac{80 + 9600}{10000} = 0.968

Precision — of everyone it flagged, how many really had it:

precision=TPTP+FP=8080+300=0.211\text{precision} = \frac{TP}{TP + FP} = \frac{80}{80 + 300} = 0.211

Recall — of everyone who had it, how many it caught:

recall=TPTP+FN=8080+20=0.800\text{recall} = \frac{TP}{TP + FN} = \frac{80}{80 + 20} = 0.800

F1 — the harmonic mean, which punishes imbalance between the two:

F1=2⋅0.211×0.8000.211+0.800=0.334F_1 = 2 \cdot \frac{0.211 \times 0.800}{0.211 + 0.800} = 0.334

Look at what those numbers say together. Accuracy of 96.8% sounds like a strong model. Precision of 21% says that four out of every five people it alarms are healthy. Recall of 80% says it catches four of every five real cases. Those are three completely different descriptions of the same model, and only one of them — the one you choose — will reach the person making the decision.

Accuracy answers "how often is it right?", which is almost never the question. The real question is "what does each kind of mistake cost?", and only precision and recall can answer it.

Choosing between precision and recall

You cannot maximise both. Lowering the decision threshold catches more true cases and also raises more false alarms; raising it does the reverse. Which way you should lean depends entirely on the relative cost of the two errors.

SituationCostly errorOptimise
Cancer screeningMissing a case (false negative)Recall
Spam filterBinning a real email (false positive)Precision
Fraud blockingFreezing a legitimate customerPrecision
Predictive maintenanceMissing a failureRecall
Search resultsAn irrelevant top hitPrecision
Python
from sklearn.metrics import classification_report, confusion_matrixproba = model.predict_proba(X_test)[:, 1]preds = (proba >= 0.30).astype(int)      # move the threshold to favour recallprint(confusion_matrix(y_test, preds))print(classification_report(y_test, preds, digits=3))

The default threshold of 0.5 is a convention, not a law. It is very rarely the right operating point for imbalanced data, and choosing it deliberately is often a bigger improvement than any change of algorithm.

ROC-AUC and its imbalanced-data trap

ROC-AUC measures ranking quality across all thresholds: the probability that a randomly chosen positive is scored above a randomly chosen negative. 0.5 is random, 1.0 is perfect.

Python
from sklearn.metrics import roc_auc_score, average_precision_scoreprint(roc_auc_score(y_test, proba))            # 0.94  -- looks superbprint(average_precision_score(y_test, proba))  # 0.31  -- the honest number

That gap is the trap. ROC-AUC's false positive rate has the enormous negative class in its denominator, so 300 false positives out of 9,900 healthy people barely moves it. Average precision — the area under the precision-recall curve — has only the positives in view and reports what you would actually experience.

MetricUse whenMisleads when
AccuracyClasses are balanced and errors cost the sameAnything rare
Precision / recall / F1Almost alwaysYou report only one of them
ROC-AUCComparing rankers on balanced dataThe positive class is rare
Average precisionRare positives — the default for imbalanceRarely

For multiclass problems, the averaging choice matters as much as the metric. average="macro" treats every class equally, so a rare class carries the same weight as a common one; average="weighted" weights by class size, letting a large class hide poor performance on small ones. If the rare classes are the point, use macro.

Regression metrics

Python
from sklearn.metrics import mean_absolute_error, mean_squared_error, r2_scoreimport numpy as npy_true = np.array([10, 20, 30, 40, 1000])y_pred = np.array([12, 18, 33, 38,  100])print(mean_absolute_error(y_true, y_pred))                  # 181.8print(np.sqrt(mean_squared_error(y_true, y_pred)))          # 402.5

Four of those five predictions are within 3 units. One is out by 900. MAE averages the errors and reports 181.8; RMSE squares them first, so the single 900 contributes 810,000 to a total of 810,021 and dominates entirely, giving 402.5.

That difference is the diagnostic. RMSE much larger than MAE means a small number of large errors; RMSE close to MAE means errors are evenly spread. Which you should report depends on whether one enormous error is worse than many small ones — for a delivery time estimate, probably yes; for a total revenue forecast, probably not.

MetricUnitsOutlier sensitivityRead as
MAESame as targetLowTypical error size
RMSESame as targetHighError size, with big misses punished
R²NoneHighFraction of variance explained
MAPEPercentageMediumRelative error — breaks if any true value is 0

R² compares your model against the simplest possible baseline: always predicting the mean. R² of 0.0 means you have matched that baseline and added nothing. R² can go negative, which means your model is worse than the mean — a result that reliably indicates a bug, usually a preprocessing step applied inconsistently between fitting and prediction.

Why one split is not enough

Python
from sklearn.model_selection import train_test_splitfrom sklearn.ensemble import RandomForestClassifierfor seed in range(5):    Xtr, Xte, ytr, yte = train_test_split(X, y, test_size=0.2, random_state=seed)    m = RandomForestClassifier(random_state=42).fit(Xtr, ytr)    print(f"seed {seed}: {m.score(Xte, yte):.3f}")# 0.847  0.881  0.859  0.902  0.864

Same data, same model, same settings — and a spread of 5.5 percentage points depending on which rows happened to land in the test set. Report the 0.902 and you are reporting luck. This also means that if you compare two models on one split and they differ by 2 points, you have learned nothing.

Cross-validation

In the examples from here on, pipe is a pipeline like the one built in the previous lesson: preprocessing, then a model step named clf.

Python
from sklearn.model_selection import cross_val_score, StratifiedKFoldcv = StratifiedKFold(n_splits=5, shuffle=True, random_state=42)scores = cross_val_score(pipe, X, y, cv=cv, scoring="f1_macro")print(f"{scores.mean():.3f} ± {scores.std():.3f}")     # 0.871 ± 0.021

The data is divided into five parts; the model is trained five times, each time holding out a different part. Every row gets to be test data exactly once, and you get a mean plus a spread. The spread is not a footnote — it is what tells you whether a 0.006 difference between two models is real.

Python
from sklearn.model_selection import cross_validateresults = cross_validate(pipe, X, y, cv=cv,                         scoring=["precision_macro", "recall_macro", "f1_macro"],                         return_train_score=True)

return_train_score=True gives you the overfitting gap for free: a training F1 of 0.99 against a validation F1 of 0.71 diagnoses the problem without any extra work.

Use StratifiedKFold for classification, plain KFold for regression, GroupKFold when rows cluster by customer or patient, and TimeSeriesSplit when the data is ordered in time — that last one trains only on earlier folds and tests on later ones, because shuffling a timeline means training on the future.

Tuning hyperparameters

Hyperparameters are the settings you choose rather than the ones the model learns: tree depth, number of neighbours, regularisation strength. Search them systematically rather than by hand.

Python
from sklearn.model_selection import GridSearchCV, RandomizedSearchCVfrom scipy.stats import randint, uniformgrid = {    "clf__n_estimators": [100, 300, 500],    "clf__max_depth": [None, 10, 20],    "clf__min_samples_leaf": [1, 5, 10],}search = GridSearchCV(pipe, grid, cv=cv, scoring="f1_macro", n_jobs=-1, refit=True)search.fit(X_train, y_train)print(search.best_params_)print(f"{search.best_score_:.3f}")final = search.best_estimator_        # already refitted on all training data

Note the double underscore in clf__max_depth. That is how you address a parameter of a named step inside a pipeline — step name, two underscores, parameter name — and it is what allows the scaler to be refitted inside every fold of every candidate configuration.

SearchFits requiredBest for
GridSearchCVEvery combination × foldsFew parameters with few sensible values
RandomizedSearchCVn_iter × folds — you chooseMany parameters, or continuous ranges
HalvingRandomSearchCVFar fewerLarge budgets; eliminates weak candidates early

The grid above is 3 × 3 × 3 = 27 combinations × 5 folds = 135 model fits. Add one more parameter with four values and it is 540. Grid search grows multiplicatively, so beyond three parameters random search finds better configurations for the same compute — most parameters barely matter, and random sampling spends its budget exploring the ones that do rather than exhaustively mapping the ones that do not.

Python
dist = {    "clf__n_estimators": randint(100, 800),    "clf__max_depth": randint(3, 30),    "clf__min_samples_leaf": randint(1, 20),    "clf__max_features": uniform(0.1, 0.9),}search = RandomizedSearchCV(pipe, dist, n_iter=60, cv=cv,                            scoring="f1_macro", n_jobs=-1, random_state=42)

The score reported by a hyperparameter search is optimistic — you selected the best of many attempts. Only an untouched test set gives an honest final number.

Diagnosing what is wrong

Training scoreValidation scoreDiagnosisFix
High (0.99)Much lower (0.71)OverfittingMore data, simpler model, regularisation, fewer features
Low (0.65)Low (0.63)UnderfittingMore capacity, better features, less regularisation
HighHighWorking — or leakingCheck for leakage before celebrating
LowHigher than trainingSomething is wrongUsually a broken split or an inconsistent transform
Python
from sklearn.model_selection import learning_curveimport matplotlib.pyplot as pltimport numpy as npsizes, train_sc, val_sc = learning_curve(    pipe, X, y, cv=cv, scoring="f1_macro",    train_sizes=np.linspace(0.1, 1.0, 8), n_jobs=-1)plt.plot(sizes, train_sc.mean(1), "o-", label="training")plt.plot(sizes, val_sc.mean(1), "o-", label="validation")plt.xlabel("Training examples"); plt.ylabel("F1"); plt.legend()

The learning curve answers a question worth real money: will more data help? If the two lines are converging and the gap is still closing, collecting more data will improve the model. If they have flattened and met at a low value, more data of the same kind will change nothing, and your effort belongs in features or in a different model family.

Feature importance, and the bias in the obvious version

Python
import pandas as pdfrom sklearn.inspection import permutation_importance# one value per column AFTER preprocessing (each one-hot level separately)built_in = pd.Series(final["clf"].feature_importances_,                     index=final[:-1].get_feature_names_out())# one value per ORIGINAL column: the whole pipeline is scoredperm = permutation_importance(final, X_test, y_test, n_repeats=20,                              random_state=42, scoring="f1_macro")honest = pd.Series(perm.importances_mean, index=X_test.columns)print(built_in.sort_values(ascending=False).head(10).round(4))print(honest.sort_values(ascending=False).round(4))

A tree's built-in importance counts how much each feature reduced impurity during training, which systematically favours features with many distinct values — a customer ID or a continuous variable offers more places to split, so it accumulates importance whether or not it generalises. Permutation importance instead shuffles one column in the test set and measures how much the score drops. If shuffling a feature does not hurt performance, the model was not really using it. The two lists are at different levels when the pipeline one-hot encodes: the built-in version scores each encoded column (cat__region_north), while permutation scores the original column (region) as a whole.

Where the two disagree sharply, believe the permutation version. A feature that ranks first by impurity and near zero by permutation is a feature the model memorised rather than learned from.

What to report when you are done

Python
from sklearn.metrics import classification_report, confusion_matrix, f1_scorefinal = search.best_estimator_preds = final.predict(X_test)          # the test set, touched once, at the endprint(confusion_matrix(y_test, preds))print(classification_report(y_test, preds, digits=3))print(f"cross-validated: {search.best_score_:.3f}")print(f"held-out test:   {f1_score(y_test, preds, average='macro'):.3f}")

Report four things, always. The metric and why you chose it, tied to what the two error types cost. The cross-validated mean and standard deviation, so the reader knows how much of any difference is noise. The held-out test score, computed once. And the confusion matrix, because it is the only output that shows the errors themselves rather than an average of them.

A model is not finished when it fits. It is finished when you can say what it gets wrong, how often, and what that costs — and when the number you are quoting came from data the model has never influenced in any way.

Check your understanding

0 of 3 answered

1.A fraud model reports 99.5% accuracy on data where 0.5% of transactions are fraud. What should you look at before trusting it?

2.A regression model has an MAE of 5 and an RMSE of 40. What does that gap tell you?

3.GridSearchCV reports a best cross-validated F1 of 0.91. Which number should go in your final report as the estimate of real performance?