Course Content
Deep Learning with TensorFlow and PyTorch
4 sections · 15 lessons
Model Evaluation and Hyperparameter Tuning
A team ships a fraud detector. The report says 99.0% accuracy. Everyone is pleased. Three weeks later the finance department asks why the system has not flagged a single fraudulent transaction.
Here is what happened. Out of 10,000 transactions, 100 were fraudulent. The model learned to output "not fraud" for everything. It is right 9,900 times out of 10,000 — 99.0% accuracy — while being completely useless at the only job it has.
Now replace it with an honest model. It flags 60 transactions; 45 of those really are fraud. It misses 55 frauds and raises 15 false alarms. Its accuracy is also about 99% (9,930 correct out of 10,000). Same headline number, radically different model. One catches 45% of fraud, the other catches none.
A single number that cannot distinguish a working model from a broken one is not a measurement. It is a comfort.
Evaluation is where most projects quietly go wrong, and the damage is usually done before training starts — in how the data was split.
Three splits, three different jobs
| Split | Typical size | Used for | How often you look at it |
|---|---|---|---|
| Training | 60–80% | Fitting the weights via gradient descent | Every step |
| Validation | 10–20% | Choosing architecture, learning rate, when to stop | Every epoch, and every time you make a decision |
| Test | 10–20% | One honest estimate of real-world performance | Exactly once, at the very end |
The reason for two held-out sets rather than one is subtle and important. Suppose you try 50 architectures and pick the one with the best validation score. That winning score is optimistically biased — with 50 attempts, some model got lucky on the particular noise in the validation set. You selected on that set, so it is no longer an unbiased estimate of anything. The test set is untouched by any decision, which is what makes its number honest.
And once you look at the test set and change something in response, it becomes a validation set. There is no undo. If you genuinely need another honest number, you need another untouched slice of data.
The leak that inflates every score you report
This is the most common serious bug in applied machine learning, and it is one line of code.
1# WRONG -- the scaler has seen the test data2from sklearn.preprocessing import StandardScaler3X_scaled = StandardScaler().fit_transform(X) # mean/std over ALL rows4X_train, X_test = train_test_split(X_scaled, test_size=0.2)The scaler computed the mean and standard deviation using the test rows. That information — where the test data sits — has now been baked into every training example. Your test score comes out a point or two too high, and you will not find out until production.
1# RIGHT -- split first, fit the scaler on training data only2X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2,3 random_state=42)4scaler = StandardScaler().fit(X_train) # statistics from training data only5X_train = scaler.transform(X_train)6X_test = scaler.transform(X_test) # apply, do not re-fitThe same principle covers imputing missing values, selecting features by correlation with the target, oversampling a minority class, and choosing a vocabulary. Every one of those must be fitted on training data and applied to the others. And if your data has a time dimension, split by time rather than at random — shuffling lets the model see next month's behaviour while predicting this month's, which is a leak no amount of cross-validation will catch.
Regression metrics
Take five house predictions, in thousands of pounds:
| True | Predicted | Error | |Error| | Error² |
|---|---|---|---|---|
| 200 | 210 | +10 | 10 | 100 |
| 250 | 240 | −10 | 10 | 100 |
| 300 | 310 | +10 | 10 | 100 |
| 350 | 330 | −20 | 20 | 400 |
| 400 | 450 | +50 | 50 | 2500 |
| Metric | Computation | Value | Reads as |
|---|---|---|---|
| MAE | 100/5 | 20.0 | "Typically off by £20,000" |
| MSE | 3200/5 | 640.0 | Not interpretable — units are pounds squared |
| RMSE | 640 | 25.3 | "Off by £25,300, with big misses weighted more" |
| MAPE | mean of ∣e∣/y | 6.1% | "Typically off by 6% of the true price" |
| R2 | 1−3200/25000 | 0.872 | "Explains 87% of the variation in price" |
RMSE exceeds MAE (25.3 versus 20.0) because of the single 50-unit miss. That gap is diagnostic: RMSE much larger than MAE means a few large errors are dominating. If they are equal, your errors are uniform in size.
R2 is worth understanding rather than reciting. It compares your model's squared error to the squared error of the dumbest possible model — one that always predicts the mean. R2=0 means you have matched that baseline; negative R2 means you have done worse than always guessing the average, which happens more often than people admit and almost always signals a bug.
MAPE has a specific trap: it explodes when true values approach zero, and it is lopsided — an under-prediction can never cost more than 100%, while an over-prediction has no ceiling, so it quietly favours models that predict low. Do not use it on data containing near-zero targets.
Classification metrics, with the fraud numbers
Everything starts from the confusion matrix. For our honest fraud model:
| Predicted fraud | Predicted legitimate | |
|---|---|---|
| Actually fraud | TP = 45 | FN = 55 |
| Actually legitimate | FP = 15 | TN = 9,885 |
| Metric | Formula | Value | The question it answers |
|---|---|---|---|
| Accuracy | (TP+TN)/total | 99.3% | How often is the model right? (Useless here) |
| Precision | TP/(TP+FP) | 45/60 = 75% | When it raises an alarm, how often is it real? |
| Recall | TP/(TP+FN) | 45/100 = 45% | Of all real fraud, how much did it catch? |
| F1 | 2PR/(P+R) | 0.563 | A single number balancing the two |
| Specificity | TN/(TN+FP) | 99.8% | Of all legitimate transactions, how many were left alone? |
Precision and recall trade against each other, and the dial that moves them is the decision threshold. Lower the threshold from 0.5 to 0.2 and the model flags more transactions: recall rises, precision falls. Raise it to 0.8 and the reverse happens. Which direction is right is a business question, not a modelling one.
| Situation | Optimise for | Why |
|---|---|---|
| Cancer screening | Recall | A missed tumour is far worse than an unnecessary follow-up scan |
| Spam filtering | Precision | A lost job offer in the spam folder costs more than a spam message in the inbox |
| Fraud with a manual review queue | Recall, capped by reviewer capacity | Catch as much as humans can actually process |
| Balanced classes, symmetric costs | Accuracy or F1 | No asymmetry to encode |
F1 is the harmonic mean, not the arithmetic mean, and that choice matters. With precision 1.0 and recall 0.0, the arithmetic mean would be a reassuring 0.5; F1 is 0.0. The harmonic mean refuses to be fooled by one good component.
ROC-AUC, and when it lies to you too
ROC-AUC measures ranking quality across all thresholds: the probability that a randomly chosen positive is scored higher than a randomly chosen negative. 0.5 is random, 1.0 is perfect. Its appeal is that it does not depend on where you set the threshold.
Its weakness shows up under heavy imbalance. Because the false-positive rate has a denominator of 9,900 legitimate transactions, adding a few hundred false positives barely moves it, and AUC stays flatteringly high. Precision-recall AUC uses precision instead, whose denominator is the small number of positive predictions, so it responds sharply to false alarms. For imbalance beyond roughly 1:100, report PR-AUC.
1from sklearn.metrics import (classification_report, confusion_matrix,2 roc_auc_score, average_precision_score,3 precision_recall_curve)4import numpy as np5import tensorflow as tf67probs = tf.sigmoid(model.predict(X_val)).numpy().ravel() # the model outputs logits8preds = (probs >= 0.5).astype(int)910print(confusion_matrix(y_val, preds))11print(classification_report(y_val, preds, digits=3))12print("ROC-AUC:", roc_auc_score(y_val, probs))13print("PR-AUC :", average_precision_score(y_val, probs))1415# Choose the threshold that maximises F1, rather than defaulting to 0.516prec, rec, thresh = precision_recall_curve(y_val, probs)17f1 = 2 * prec * rec / (prec + rec + 1e-12)18print("best threshold:", thresh[np.argmax(f1[:-1])])That last block is worth making a habit. The 0.5 threshold is a convention with no claim to optimality; choosing it on the validation set is free performance.
Hyperparameters versus parameters
| Parameters | Hyperparameters | |
|---|---|---|
| Examples | Weights, biases | Learning rate, layer widths, dropout rate, batch size |
| How they are set | Learned by gradient descent | Chosen by you, before training |
| How many | Millions | Usually under twenty |
| Selected using | Training data | Validation data |
Their influence is not equal, and knowing the ordering saves enormous time:
- Learning rate. Wrong by a factor of ten and nothing else you do matters.
- Architecture size — depth and width. Sets capacity.
- Regularisation strength — dropout rate, weight decay. Only matters once you are overfitting.
- Batch size. Mostly a throughput decision; interacts with the learning rate.
- Optimiser betas, epsilon. Almost never worth touching.
Searching the space
Why random search beats grid search
Grid search tries every combination on a regular lattice. With three hyperparameters and five values each, that is 125 training runs — and it explores exactly five distinct values of each hyperparameter. Random search with the same 125-run budget explores 125 distinct values of each.
This matters because hyperparameter importance is wildly uneven. If learning rate dominates and the other two barely matter, grid search has spent 125 runs learning about five learning rates. Random search has spent 125 runs learning about 125 of them. Same cost, twenty-five times the resolution on the parameter that counts.
1import numpy as np23def sample_config(rng):4 return {5 # Sample the learning rate LOG-uniformly: the meaningful difference6 # is between 1e-4 and 1e-3, not between 0.05 and 0.06.7 "lr": 10 ** rng.uniform(-4.5, -2.0),8 "width": int(rng.choice([64, 128, 256, 512])),9 "depth": int(rng.integers(2, 6)),10 "dropout": rng.uniform(0.0, 0.5),11 "weight_decay": 10 ** rng.uniform(-6, -2),12 }1314rng = np.random.default_rng(0)15best = (float("inf"), None)16for trial in range(40):17 cfg = sample_config(rng)18 val_loss = train_and_evaluate(cfg) # your training function19 if val_loss < best[0]:20 best = (val_loss, cfg)21 print(f"trial {trial:3d} val={val_loss:.4f} {cfg}")22print("best:", best)The log-uniform sampling of the learning rate is not a detail. Sampling uniformly between 0.0001 and 0.01 puts 90% of your samples above 0.001, so you would barely explore the small end of the range. Anything spanning orders of magnitude — learning rate, weight decay, regularisation strength — should be sampled in log space.
Bayesian search, when each run is expensive
Random search treats every trial as independent. Bayesian optimisation builds a model of "which regions of the space produce good scores" and samples where that model is most uncertain or most promising. It typically finds a comparable configuration in a third to a half the trials — worth it when a single run takes hours.
1import optuna23def objective(trial):4 cfg = {5 "lr": trial.suggest_float("lr", 1e-5, 1e-2, log=True),6 "width": trial.suggest_categorical("width", [64, 128, 256, 512]),7 "depth": trial.suggest_int("depth", 2, 5),8 "dropout": trial.suggest_float("dropout", 0.0, 0.5),9 }10 return train_and_evaluate(cfg)1112study = optuna.create_study(13 direction="minimize",14 pruner=optuna.pruners.MedianPruner(), # abandon hopeless trials early15)16study.optimize(objective, n_trials=50)17print(study.best_params, study.best_value)The pruner is where the real saving comes from. It stops a trial part-way when its intermediate scores are worse than the median of previous trials at the same epoch, so obviously bad configurations cost a fraction of a full run.
Cross-validation, and when it is worth the cost
With a small dataset, a single validation split is a lottery: 200 held-out examples give an estimate with several points of noise, so you cannot tell a real 1% improvement from luck. k-fold cross-validation splits the data k ways, trains k times, and averages — giving both a better estimate and a standard deviation.
1from sklearn.model_selection import StratifiedKFold2import numpy as np34scores = []5for fold, (tr, va) in enumerate(StratifiedKFold(5, shuffle=True,6 random_state=42).split(X, y)):7 model = build_model() # fresh model every fold8 scaler = StandardScaler().fit(X[tr]) # fresh scaler every fold9 model.fit(scaler.transform(X[tr]), y[tr])10 scores.append(model.score(scaler.transform(X[va]), y[va]))1112print(f"{np.mean(scores):.4f} +/- {np.std(scores):.4f}")Two rules. Rebuild the model inside the loop — reusing a fitted model means later folds are evaluated on data the model trained on. And use stratified folds for classification, so each fold preserves the class balance; with 100 positives out of 10,000, plain random folds can easily produce a fold with almost none.
The honest cost is that cross-validation is k times the compute. On a deep model that trains for six hours, five-fold cross-validation is a day and a quarter. It is standard practice below roughly 10,000 examples and rare above 100,000, where a single split is already large enough to be stable.
Learning curves tell you what to fix
Plot training and validation loss against epoch. The shape of that pair of curves narrows the diagnosis faster than any other single tool.
| Training loss | Validation loss | Diagnosis | What to do |
|---|---|---|---|
| High, flat | High, flat | Underfitting — not enough capacity or training | Bigger model, longer training, higher learning rate |
| Very low | Much higher and rising | Overfitting | More data, dropout, weight decay, early stopping |
| Low | Low, close to training | Healthy | Stop tuning; consider whether more capacity helps |
| Falling | Falling, still going down at the end | Under-trained | Train for more epochs |
| Erratic, spiky | Erratic | Learning rate too high, or batch too small | Lower the LR; add a schedule |
| Falls fast then flat | Flat above training | Learning rate too high for the current region | Add a decay schedule |
There is a second learning curve worth plotting: final score against training set size. If the curve is still rising steeply when you run out of data, more data is the highest-value investment available. If it has flattened, collecting more data will not help and the bottleneck is the model or the features. That single plot has settled a great many arguments about where to spend the next month.
What to actually do on a project
Split the data first, before any preprocessing, and if there is a time dimension split by time. Fit every transformation on the training portion only. Choose your primary metric from the cost of each kind of error, and write down what that cost is — "a false negative costs us £400, a false positive costs us 10 minutes of an analyst's time" turns a modelling argument into an arithmetic one.
Then get a baseline before you tune anything: the majority-class predictor for classification, the mean predictor for regression, and if you can manage it, a logistic regression or gradient-boosted tree. If your neural network cannot beat gradient boosting on tabular data, the problem is the choice of model, and no hyperparameter search will fix it.
Tune the learning rate first and alone. Then run a random search of thirty to fifty trials over the remaining parameters, sampling anything scale-free in log space. Track every trial's configuration and score; the log is more valuable than the winner, because it tells you which parameters mattered.
And touch the test set exactly once, when you are finished. Report that number, along with the validation number, and be explicit about the gap between them. That gap is a measure of how much you overfitted your own search process, and reporting it honestly is what separates a result from a claim.