Deep Learning with TensorFlow and PyTorch

Model Evaluation and Hyperparameter Tuning


A team ships a fraud detector. The report says 99.0% accuracy. Everyone is pleased. Three weeks later the finance department asks why the system has not flagged a single fraudulent transaction.

Here is what happened. Out of 10,000 transactions, 100 were fraudulent. The model learned to output "not fraud" for everything. It is right 9,900 times out of 10,000 — 99.0% accuracy — while being completely useless at the only job it has.

Now replace it with an honest model. It flags 60 transactions; 45 of those really are fraud. It misses 55 frauds and raises 15 false alarms. Its accuracy is also about 99% (9,930 correct out of 10,000). Same headline number, radically different model. One catches 45% of fraud, the other catches none.

A single number that cannot distinguish a working model from a broken one is not a measurement. It is a comfort.

Evaluation is where most projects quietly go wrong, and the damage is usually done before training starts — in how the data was split.

99% accurate, and it has never caught a fraud01,000099,000flaggednot flaggedactually fraudactually cleanRecall on fraud is 0 of 1,000; accuracy is 99% because the majority class carries it.
Accuracy averages over a class ratio you did not choose — on 1% fraud, predicting "clean" always scores 99%.

Three splits, three different jobs

SplitTypical sizeUsed forHow often you look at it
Training60–80%Fitting the weights via gradient descentEvery step
Validation10–20%Choosing architecture, learning rate, when to stopEvery epoch, and every time you make a decision
Test10–20%One honest estimate of real-world performanceExactly once, at the very end

The reason for two held-out sets rather than one is subtle and important. Suppose you try 50 architectures and pick the one with the best validation score. That winning score is optimistically biased — with 50 attempts, some model got lucky on the particular noise in the validation set. You selected on that set, so it is no longer an unbiased estimate of anything. The test set is untouched by any decision, which is what makes its number honest.

And once you look at the test set and change something in response, it becomes a validation set. There is no undo. If you genuinely need another honest number, you need another untouched slice of data.

The leak that inflates every score you report

This is the most common serious bug in applied machine learning, and it is one line of code.

Python
# WRONG -- the scaler has seen the test datafrom sklearn.preprocessing import StandardScalerX_scaled = StandardScaler().fit_transform(X)          # mean/std over ALL rowsX_train, X_test = train_test_split(X_scaled, test_size=0.2)

The scaler computed the mean and standard deviation using the test rows. That information — where the test data sits — has now been baked into every training example. Your test score comes out a point or two too high, and you will not find out until production.

Python
# RIGHT -- split first, fit the scaler on training data onlyX_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2,                                                    random_state=42)scaler = StandardScaler().fit(X_train)     # statistics from training data onlyX_train = scaler.transform(X_train)X_test  = scaler.transform(X_test)         # apply, do not re-fit

The same principle covers imputing missing values, selecting features by correlation with the target, oversampling a minority class, and choosing a vocabulary. Every one of those must be fitted on training data and applied to the others. And if your data has a time dimension, split by time rather than at random — shuffling lets the model see next month's behaviour while predicting this month's, which is a leak no amount of cross-validation will catch.

Regression metrics

Take five house predictions, in thousands of pounds:

TruePredictedError|Error|Error²
200210+1010100
250240−1010100
300310+1010100
350330−2020400
400450+50502500
MetricComputationValueReads as
MAE100/5100 / 520.0"Typically off by £20,000"
MSE3200/53200 / 5640.0Not interpretable — units are pounds squared
RMSE640\sqrt{640}25.3"Off by £25,300, with big misses weighted more"
MAPEmean of ∣e∣/y|e|/y6.1%"Typically off by 6% of the true price"
R2R^21−3200/250001 - 3200/250000.872"Explains 87% of the variation in price"

RMSE exceeds MAE (25.3 versus 20.0) because of the single 50-unit miss. That gap is diagnostic: RMSE much larger than MAE means a few large errors are dominating. If they are equal, your errors are uniform in size.

R2R^2 is worth understanding rather than reciting. It compares your model's squared error to the squared error of the dumbest possible model — one that always predicts the mean. R2=0R^2 = 0 means you have matched that baseline; negative R2R^2 means you have done worse than always guessing the average, which happens more often than people admit and almost always signals a bug.

MAPE has a specific trap: it explodes when true values approach zero, and it is lopsided — an under-prediction can never cost more than 100%, while an over-prediction has no ceiling, so it quietly favours models that predict low. Do not use it on data containing near-zero targets.

Classification metrics, with the fraud numbers

Everything starts from the confusion matrix. For our honest fraud model:

Predicted fraudPredicted legitimate
Actually fraudTP = 45FN = 55
Actually legitimateFP = 15TN = 9,885
MetricFormulaValueThe question it answers
Accuracy(TP+TN)/total(TP+TN)/\text{total}99.3%How often is the model right? (Useless here)
PrecisionTP/(TP+FP)TP/(TP+FP)45/60 = 75%When it raises an alarm, how often is it real?
RecallTP/(TP+FN)TP/(TP+FN)45/100 = 45%Of all real fraud, how much did it catch?
F12PR/(P+R)2PR/(P+R)0.563A single number balancing the two
SpecificityTN/(TN+FP)TN/(TN+FP)99.8%Of all legitimate transactions, how many were left alone?

Precision and recall trade against each other, and the dial that moves them is the decision threshold. Lower the threshold from 0.5 to 0.2 and the model flags more transactions: recall rises, precision falls. Raise it to 0.8 and the reverse happens. Which direction is right is a business question, not a modelling one.

SituationOptimise forWhy
Cancer screeningRecallA missed tumour is far worse than an unnecessary follow-up scan
Spam filteringPrecisionA lost job offer in the spam folder costs more than a spam message in the inbox
Fraud with a manual review queueRecall, capped by reviewer capacityCatch as much as humans can actually process
Balanced classes, symmetric costsAccuracy or F1No asymmetry to encode

F1 is the harmonic mean, not the arithmetic mean, and that choice matters. With precision 1.0 and recall 0.0, the arithmetic mean would be a reassuring 0.5; F1 is 0.0. The harmonic mean refuses to be fooled by one good component.

ROC-AUC, and when it lies to you too

ROC-AUC measures ranking quality across all thresholds: the probability that a randomly chosen positive is scored higher than a randomly chosen negative. 0.5 is random, 1.0 is perfect. Its appeal is that it does not depend on where you set the threshold.

Its weakness shows up under heavy imbalance. Because the false-positive rate has a denominator of 9,900 legitimate transactions, adding a few hundred false positives barely moves it, and AUC stays flatteringly high. Precision-recall AUC uses precision instead, whose denominator is the small number of positive predictions, so it responds sharply to false alarms. For imbalance beyond roughly 1:100, report PR-AUC.

Python
from sklearn.metrics import (classification_report, confusion_matrix,                             roc_auc_score, average_precision_score,                             precision_recall_curve)import numpy as npimport tensorflow as tfprobs = tf.sigmoid(model.predict(X_val)).numpy().ravel()   # the model outputs logitspreds = (probs >= 0.5).astype(int)print(confusion_matrix(y_val, preds))print(classification_report(y_val, preds, digits=3))print("ROC-AUC:", roc_auc_score(y_val, probs))print("PR-AUC :", average_precision_score(y_val, probs))# Choose the threshold that maximises F1, rather than defaulting to 0.5prec, rec, thresh = precision_recall_curve(y_val, probs)f1 = 2 * prec * rec / (prec + rec + 1e-12)print("best threshold:", thresh[np.argmax(f1[:-1])])

That last block is worth making a habit. The 0.5 threshold is a convention with no claim to optimality; choosing it on the validation set is free performance.

Hyperparameters versus parameters

ParametersHyperparameters
ExamplesWeights, biasesLearning rate, layer widths, dropout rate, batch size
How they are setLearned by gradient descentChosen by you, before training
How manyMillionsUsually under twenty
Selected usingTraining dataValidation data

Their influence is not equal, and knowing the ordering saves enormous time:

  1. Learning rate. Wrong by a factor of ten and nothing else you do matters.
  2. Architecture size — depth and width. Sets capacity.
  3. Regularisation strength — dropout rate, weight decay. Only matters once you are overfitting.
  4. Batch size. Mostly a throughput decision; interacts with the learning rate.
  5. Optimiser betas, epsilon. Almost never worth touching.

Searching the space

Why random search beats grid search

Grid search tries every combination on a regular lattice. With three hyperparameters and five values each, that is 125 training runs — and it explores exactly five distinct values of each hyperparameter. Random search with the same 125-run budget explores 125 distinct values of each.

This matters because hyperparameter importance is wildly uneven. If learning rate dominates and the other two barely matter, grid search has spent 125 runs learning about five learning rates. Random search has spent 125 runs learning about 125 of them. Same cost, twenty-five times the resolution on the parameter that counts.

Python
import numpy as npdef sample_config(rng):    return {        # Sample the learning rate LOG-uniformly: the meaningful difference        # is between 1e-4 and 1e-3, not between 0.05 and 0.06.        "lr": 10 ** rng.uniform(-4.5, -2.0),        "width": int(rng.choice([64, 128, 256, 512])),        "depth": int(rng.integers(2, 6)),        "dropout": rng.uniform(0.0, 0.5),        "weight_decay": 10 ** rng.uniform(-6, -2),    }rng = np.random.default_rng(0)best = (float("inf"), None)for trial in range(40):    cfg = sample_config(rng)    val_loss = train_and_evaluate(cfg)          # your training function    if val_loss < best[0]:        best = (val_loss, cfg)    print(f"trial {trial:3d}  val={val_loss:.4f}  {cfg}")print("best:", best)

The log-uniform sampling of the learning rate is not a detail. Sampling uniformly between 0.0001 and 0.01 puts 90% of your samples above 0.001, so you would barely explore the small end of the range. Anything spanning orders of magnitude — learning rate, weight decay, regularisation strength — should be sampled in log space.

Bayesian search, when each run is expensive

Random search treats every trial as independent. Bayesian optimisation builds a model of "which regions of the space produce good scores" and samples where that model is most uncertain or most promising. It typically finds a comparable configuration in a third to a half the trials — worth it when a single run takes hours.

Python
import optunadef objective(trial):    cfg = {        "lr": trial.suggest_float("lr", 1e-5, 1e-2, log=True),        "width": trial.suggest_categorical("width", [64, 128, 256, 512]),        "depth": trial.suggest_int("depth", 2, 5),        "dropout": trial.suggest_float("dropout", 0.0, 0.5),    }    return train_and_evaluate(cfg)study = optuna.create_study(    direction="minimize",    pruner=optuna.pruners.MedianPruner(),   # abandon hopeless trials early)study.optimize(objective, n_trials=50)print(study.best_params, study.best_value)

The pruner is where the real saving comes from. It stops a trial part-way when its intermediate scores are worse than the median of previous trials at the same epoch, so obviously bad configurations cost a fraction of a full run.

Cross-validation, and when it is worth the cost

With a small dataset, a single validation split is a lottery: 200 held-out examples give an estimate with several points of noise, so you cannot tell a real 1% improvement from luck. kk-fold cross-validation splits the data kk ways, trains kk times, and averages — giving both a better estimate and a standard deviation.

Python
from sklearn.model_selection import StratifiedKFoldimport numpy as npscores = []for fold, (tr, va) in enumerate(StratifiedKFold(5, shuffle=True,                                                random_state=42).split(X, y)):    model = build_model()                      # fresh model every fold    scaler = StandardScaler().fit(X[tr])       # fresh scaler every fold    model.fit(scaler.transform(X[tr]), y[tr])    scores.append(model.score(scaler.transform(X[va]), y[va]))print(f"{np.mean(scores):.4f} +/- {np.std(scores):.4f}")

Two rules. Rebuild the model inside the loop — reusing a fitted model means later folds are evaluated on data the model trained on. And use stratified folds for classification, so each fold preserves the class balance; with 100 positives out of 10,000, plain random folds can easily produce a fold with almost none.

The honest cost is that cross-validation is kk times the compute. On a deep model that trains for six hours, five-fold cross-validation is a day and a quarter. It is standard practice below roughly 10,000 examples and rare above 100,000, where a single split is already large enough to be stable.

Learning curves tell you what to fix

Plot training and validation loss against epoch. The shape of that pair of curves narrows the diagnosis faster than any other single tool.

Training lossValidation lossDiagnosisWhat to do
High, flatHigh, flatUnderfitting — not enough capacity or trainingBigger model, longer training, higher learning rate
Very lowMuch higher and risingOverfittingMore data, dropout, weight decay, early stopping
LowLow, close to trainingHealthyStop tuning; consider whether more capacity helps
FallingFalling, still going down at the endUnder-trainedTrain for more epochs
Erratic, spikyErraticLearning rate too high, or batch too smallLower the LR; add a schedule
Falls fast then flatFlat above trainingLearning rate too high for the current regionAdd a decay schedule

There is a second learning curve worth plotting: final score against training set size. If the curve is still rising steeply when you run out of data, more data is the highest-value investment available. If it has flattened, collecting more data will not help and the bottleneck is the model or the features. That single plot has settled a great many arguments about where to spend the next month.

What to actually do on a project

Split the data first, before any preprocessing, and if there is a time dimension split by time. Fit every transformation on the training portion only. Choose your primary metric from the cost of each kind of error, and write down what that cost is — "a false negative costs us £400, a false positive costs us 10 minutes of an analyst's time" turns a modelling argument into an arithmetic one.

Then get a baseline before you tune anything: the majority-class predictor for classification, the mean predictor for regression, and if you can manage it, a logistic regression or gradient-boosted tree. If your neural network cannot beat gradient boosting on tabular data, the problem is the choice of model, and no hyperparameter search will fix it.

Tune the learning rate first and alone. Then run a random search of thirty to fifty trials over the remaining parameters, sampling anything scale-free in log space. Track every trial's configuration and score; the log is more valuable than the winner, because it tells you which parameters mattered.

And touch the test set exactly once, when you are finished. Report that number, along with the validation number, and be explicit about the gap between them. That gap is a measure of how much you overfitted your own search process, and reporting it honestly is what separates a result from a claim.