Machine Learning Foundations

Course Content

Machine Learning Foundations

14 sections · 70 lessons

What is the difference between model parameters and hyperparameters?


What you need to know

A simple test: if the training algorithm sets it, it is a parameter. If you had to type it in, it is a hyperparameter.

Parameters

  • Learned from the data by the training algorithm
  • Examples: weights, biases, tree split thresholds
  • Can number in the millions or billions
  • Saved inside the model file

Hyperparameters

  • Chosen by you before training starts
  • Examples: learning rate, max_depth, n_estimators, C, k
  • Usually a handful
  • Saved in your config and experiment log

Seeing both in scikit-learn

Python
import numpy as npfrom sklearn.linear_model import LogisticRegressionrng = np.random.default_rng(0)X = rng.normal(size=(200, 3))y = (X[:, 0] + 0.5 * X[:, 1] > 0).astype(int)model = LogisticRegression(C=0.5, max_iter=200)      # hyperparameters: chosen by youmodel.fit(X, y)print("hyperparameters:", {k: model.get_params()[k] for k in ["C", "max_iter", "l1_ratio"]})print("parameters:     ", model.coef_.round(2), model.intercept_.round(2))   # learned from data
Text
hyperparameters: {'C': 0.5, 'max_iter': 200, 'l1_ratio': 0.0}parameters:      [[3.42 1.69 0.1 ]] [-0.08]

The hyperparameters are exactly what we passed in (plus defaults). The parameters were found by fit(). Notice they make sense: the data was built from feature 0 and half of feature 1, and the model gave feature 0 about twice the weight of feature 1, and feature 2 almost none. scikit-learn marks learned values with a trailing underscore — coef_, intercept_, tree_ — which is a handy way to tell them apart.

Common hyperparameters by model

ModelKey hyperparameters
Logistic / linear regressionC or alpha (regularisation strength), L1/L2 mix
Decision treemax_depth, min_samples_leaf
Random forestn_estimators, max_depth, max_features
Gradient boostinglearning_rate, number of trees, max_leaf_nodes, min_samples_leaf
k-NNk (n_neighbors), distance metric
Neural networklearning rate, batch size, epochs, layers and width, dropout

Grey areas worth knowing

  • The number of boosting rounds with early stopping — you set a maximum and a patience, but the actual number is chosen by the training process from validation loss.
  • Preprocessing choices — which scaler, which encoding, how many PCA components — are hyperparameters of the whole pipeline and should be tuned the same way.
  • The decision threshold — chosen after training, on validation data, so it behaves like a hyperparameter too.

A real-life example

A food-delivery company retrains its ETA model every week. One Monday the new model is noticeably worse, although the code did not change. The team compares experiment logs and finds that a colleague had changed learning_rate from 0.05 to 0.2 in a notebook and the value leaked into the config. Because hyperparameters were logged with every run, alongside the data snapshot, they found the cause in ten minutes. The parameters — millions of tree split values — were never meant to be compared by hand; the hyperparameters were the small, human-controlled part that explained the change.

Follow-up questions to expect

  • "Is the number of layers in a neural network a hyperparameter?" — Yes. Anything about the architecture that you choose before training is a hyperparameter.
  • "Can hyperparameters be learned?" — Not by the normal training step, but they can be searched automatically with grid search, random search or Bayesian optimisation, all using a validation score.
  • "Why should the test set not be used to pick hyperparameters?" — Picking by test score makes the test score optimistic; it is no longer an unbiased estimate of performance on new data.