LLMs Deep Dive

Course Content

LLMs Deep Dive

10 sections · 40 lessons

What is a hyperparameter?


What you need to know

Parameters

  • Learned from data by gradient descent
  • Billions of numbers (the weights)
  • Saved in the model file
  • Example: the embedding row for " refund"

Hyperparameters

  • Chosen by the engineer before or around training
  • A handful of settings
  • Saved in the config
  • Example: learning rate 1e-4, LoRA rank 16

Three groups

GroupExamplesWho tunes it
TrainingLearning rate, schedule, warmup steps, batch size, epochs, weight decay, dropout, gradient clipping, LoRA rank/alpha/targetsML engineer fine-tuning a model
ArchitectureLayers, hidden size, heads, KV heads, vocabulary size, number of expertsModel builders, rarely app teams
InferenceTemperature, top-p, top-k, max output tokens, stop sequences, reasoning effort, penaltiesEvery application engineer

Why the learning rate matters most

Too high and the loss jumps around or diverges to NaN; too low and training crawls or stops at a poor point. Typical starting points: about 1e-4 to 2e-4 for LoRA, about 1e-5 for full fine-tuning, with a warmup of a few percent of steps and a cosine or linear decay.

How to search

  • Grid search — try every combination; wasteful beyond two or three settings.
  • Random search — sample combinations; usually finds good values faster than grid.
  • Bayesian optimisation — uses earlier results to pick the next trial (tools such as Optuna).
Python
import optunadef objective(trial):    lr = trial.suggest_float("lr", 1e-5, 5e-4, log=True)    rank = trial.suggest_categorical("rank", [8, 16, 32])    epochs = trial.suggest_int("epochs", 1, 3)    return train_and_eval(lr=lr, rank=rank, epochs=epochs)  # your function: returns validation lossstudy = optuna.create_study(direction="minimize")study.optimize(objective, n_trials=20)print(study.best_params)

Optuna samples a learning rate on a log scale, a LoRA rank and an epoch count, calls your training function, and steers later trials towards the lowest validation loss. The test set is used once, at the end.

A real-life example

A legal-tech team fine-tunes a summariser with LoRA and first uses a learning rate of 1e-3 copied from a blog post. Loss drops fast, then spikes, and outputs become repetitive. They run 12 trials: learning rates of 5e-5, 1e-4 and 2e-4, ranks of 8 and 16, and 1 or 2 epochs. The best on validation is 1e-4, rank 16, 2 epochs.

At inference they tune separately: temperature 0.2 for summaries, a maximum of 800 output tokens so answers do not ramble, and a stop sequence after the final "Risk flag" line. For the clause-analysis route running on a reasoning model, they set effort to "medium", because "high" doubled latency without improving their evaluation scores.

Follow-up questions to expect

  • "Why not tune on the test set?" — The test score would then be optimistic, since you picked settings that happen to suit it; it would no longer measure performance on unseen data.
  • "What is learning-rate warmup?" — Starting with a very small learning rate and raising it over the first steps, so early, noisy gradients do not push the weights too far.
  • "Is temperature a hyperparameter?" — It is an inference setting rather than a training one, but people commonly call it a hyperparameter; say which kind you mean.