Course Content
Deep Learning Essentials
13 sections · 61 lessons
What are hyperparameters, and which ones does a neural network have?
What you need to know
Parameters versus hyperparameters
The weights of a ResNet-50 are parameters: 25 million numbers learned from data. "Use ResNet-50, AdamW, learning rate 3e-4, batch size 64, 30 epochs" is a list of hyperparameters.
The main groups
| Group | Hyperparameters |
|---|---|
| Architecture | Number of layers, units per layer, filters, kernel size, stride, padding, activation, normalisation type |
| Optimisation | Learning rate, optimiser (SGD, AdamW), momentum and betas, schedule and warmup, batch size |
| Training length | Number of epochs or steps, early-stopping patience |
| Regularisation | Dropout rate, weight decay, label smoothing, augmentation strength, gradient-clipping threshold |
| Data | Input size or sequence length, class weights, train and validation split |
| Other | Initialisation scheme, random seed, mixed precision |
What to tune first
Not all hyperparameters matter equally. A practical order:
- Learning rate — the biggest effect by far. Find a good range with a learning-rate range test.
- Batch size and schedule — pick batch size for your GPU, then retune the learning rate; add warmup and decay.
- Regularisation — weight decay, dropout, augmentation strength, once you can see overfitting.
- Architecture size — layers and width, only if the model clearly underfits or overfits.
How to search
- Grid search tries every combination. It wastes trials, because most grid points vary hyperparameters that matter little.
- Random search samples combinations at random. For the same number of trials it usually finds better settings, because it tries more distinct values of the few hyperparameters that matter.
- Bayesian optimisation, for example Optuna's default sampler, uses the results so far to choose promising settings next, and can stop poor trials early.
1import optuna23def objective(trial):4 lr = trial.suggest_float("lr", 1e-5, 1e-2, log=True)5 dropout = trial.suggest_float("dropout", 0.0, 0.5)6 batch_size = trial.suggest_categorical("batch_size", [32, 64, 128])7 return train_and_validate(lr, dropout, batch_size) # returns validation loss89study = optuna.create_study(direction="minimize")10study.optimize(objective, n_trials=50)11print(study.best_params)log=True samples the learning rate evenly across orders of magnitude, so 1e-5 to 1e-4 gets as many trials as 1e-3 to 1e-2. train_and_validate is your own function that trains with those settings and returns the validation loss.
A real-life example
A retail chain builds an LSTM to forecast daily sales for 600 stores. The first model, with default settings, is only slightly better than "same as last week". The team lists the hyperparameters that plausibly matter: learning rate, the look-back window (28, 56 or 90 days), hidden size (32 to 256), number of LSTM layers (1 or 2), dropout and batch size.
They run 60 Optuna trials, each on a time-based split: train on data up to March, validate on April to June, never touching the July to September test period. The best trial uses a 90-day window, which captures monthly salary-day patterns, hidden size 128, 2 layers, dropout 0.2 and learning rate 8e-4. Its forecast error on the untouched test period is 18% lower than the default model's. The learning rate and look-back window explained most of the gain; hidden size barely mattered.
Follow-up questions to expect
- "Why not tune on the test set?" — The test set would then have influenced your choices, so its score is optimistic and no longer an honest estimate of performance on new data.
- "Is the number of epochs a hyperparameter?" — Yes, though it is usually handled by early stopping rather than tuned directly.
- "Why is the learning rate searched on a log scale?" — Its useful values span several orders of magnitude, and the difference between 1e-4 and 1e-3 matters as much as between 1e-3 and 1e-2.