Course Content
Machine Learning Foundations
14 sections · 70 lessons
What is the difference between model parameters and hyperparameters?
What you need to know
A simple test: if the training algorithm sets it, it is a parameter. If you had to type it in, it is a hyperparameter.
Parameters
- Learned from the data by the training algorithm
- Examples: weights, biases, tree split thresholds
- Can number in the millions or billions
- Saved inside the model file
Hyperparameters
- Chosen by you before training starts
- Examples: learning rate,
max_depth,n_estimators,C, k - Usually a handful
- Saved in your config and experiment log
Seeing both in scikit-learn
1import numpy as np2from sklearn.linear_model import LogisticRegression34rng = np.random.default_rng(0)5X = rng.normal(size=(200, 3))6y = (X[:, 0] + 0.5 * X[:, 1] > 0).astype(int)78model = LogisticRegression(C=0.5, max_iter=200) # hyperparameters: chosen by you9model.fit(X, y)10print("hyperparameters:", {k: model.get_params()[k] for k in ["C", "max_iter", "l1_ratio"]})11print("parameters: ", model.coef_.round(2), model.intercept_.round(2)) # learned from datahyperparameters: {'C': 0.5, 'max_iter': 200, 'l1_ratio': 0.0}parameters: [[3.42 1.69 0.1 ]] [-0.08]The hyperparameters are exactly what we passed in (plus defaults). The parameters were found by fit(). Notice they make sense: the data was built from feature 0 and half of feature 1, and the model gave feature 0 about twice the weight of feature 1, and feature 2 almost none. scikit-learn marks learned values with a trailing underscore — coef_, intercept_, tree_ — which is a handy way to tell them apart.
Common hyperparameters by model
| Model | Key hyperparameters |
|---|---|
| Logistic / linear regression | C or alpha (regularisation strength), L1/L2 mix |
| Decision tree | max_depth, min_samples_leaf |
| Random forest | n_estimators, max_depth, max_features |
| Gradient boosting | learning_rate, number of trees, max_leaf_nodes, min_samples_leaf |
| k-NN | k (n_neighbors), distance metric |
| Neural network | learning rate, batch size, epochs, layers and width, dropout |
Grey areas worth knowing
- The number of boosting rounds with early stopping — you set a maximum and a patience, but the actual number is chosen by the training process from validation loss.
- Preprocessing choices — which scaler, which encoding, how many PCA components — are hyperparameters of the whole pipeline and should be tuned the same way.
- The decision threshold — chosen after training, on validation data, so it behaves like a hyperparameter too.
A real-life example
A food-delivery company retrains its ETA model every week. One Monday the new model is noticeably worse, although the code did not change. The team compares experiment logs and finds that a colleague had changed learning_rate from 0.05 to 0.2 in a notebook and the value leaked into the config. Because hyperparameters were logged with every run, alongside the data snapshot, they found the cause in ten minutes. The parameters — millions of tree split values — were never meant to be compared by hand; the hyperparameters were the small, human-controlled part that explained the change.
Follow-up questions to expect
- "Is the number of layers in a neural network a hyperparameter?" — Yes. Anything about the architecture that you choose before training is a hyperparameter.
- "Can hyperparameters be learned?" — Not by the normal training step, but they can be searched automatically with grid search, random search or Bayesian optimisation, all using a validation score.
- "Why should the test set not be used to pick hyperparameters?" — Picking by test score makes the test score optimistic; it is no longer an unbiased estimate of performance on new data.