Course Content
LLMs Deep Dive
10 sections · 40 lessons
What is a hyperparameter?
What you need to know
Parameters
- Learned from data by gradient descent
- Billions of numbers (the weights)
- Saved in the model file
- Example: the embedding row for " refund"
Hyperparameters
- Chosen by the engineer before or around training
- A handful of settings
- Saved in the config
- Example: learning rate 1e-4, LoRA rank 16
Three groups
| Group | Examples | Who tunes it |
|---|---|---|
| Training | Learning rate, schedule, warmup steps, batch size, epochs, weight decay, dropout, gradient clipping, LoRA rank/alpha/targets | ML engineer fine-tuning a model |
| Architecture | Layers, hidden size, heads, KV heads, vocabulary size, number of experts | Model builders, rarely app teams |
| Inference | Temperature, top-p, top-k, max output tokens, stop sequences, reasoning effort, penalties | Every application engineer |
Why the learning rate matters most
Too high and the loss jumps around or diverges to NaN; too low and training crawls or stops at a poor point. Typical starting points: about 1e-4 to 2e-4 for LoRA, about 1e-5 for full fine-tuning, with a warmup of a few percent of steps and a cosine or linear decay.
How to search
- Grid search — try every combination; wasteful beyond two or three settings.
- Random search — sample combinations; usually finds good values faster than grid.
- Bayesian optimisation — uses earlier results to pick the next trial (tools such as Optuna).
1import optuna23def objective(trial):4 lr = trial.suggest_float("lr", 1e-5, 5e-4, log=True)5 rank = trial.suggest_categorical("rank", [8, 16, 32])6 epochs = trial.suggest_int("epochs", 1, 3)7 return train_and_eval(lr=lr, rank=rank, epochs=epochs) # your function: returns validation loss89study = optuna.create_study(direction="minimize")10study.optimize(objective, n_trials=20)11print(study.best_params)Optuna samples a learning rate on a log scale, a LoRA rank and an epoch count, calls your training function, and steers later trials towards the lowest validation loss. The test set is used once, at the end.
A real-life example
A legal-tech team fine-tunes a summariser with LoRA and first uses a learning rate of 1e-3 copied from a blog post. Loss drops fast, then spikes, and outputs become repetitive. They run 12 trials: learning rates of 5e-5, 1e-4 and 2e-4, ranks of 8 and 16, and 1 or 2 epochs. The best on validation is 1e-4, rank 16, 2 epochs.
At inference they tune separately: temperature 0.2 for summaries, a maximum of 800 output tokens so answers do not ramble, and a stop sequence after the final "Risk flag" line. For the clause-analysis route running on a reasoning model, they set effort to "medium", because "high" doubled latency without improving their evaluation scores.
Follow-up questions to expect
- "Why not tune on the test set?" — The test score would then be optimistic, since you picked settings that happen to suit it; it would no longer measure performance on unseen data.
- "What is learning-rate warmup?" — Starting with a very small learning rate and raising it over the first steps, so early, noisy gradients do not push the weights too far.
- "Is temperature a hyperparameter?" — It is an inference setting rather than a training one, but people commonly call it a hyperparameter; say which kind you mean.