Machine Learning Foundations

Course Content

Machine Learning Foundations

14 sections · 70 lessons

Why does gradient descent converge faster with scaled features?


Same data, same algorithm, different unitsRaw area and bedrooms• Stretched valley in the loss• lr 1e-6 diverges at step 458• lr 1e-7: MSE 122.6 after 100,000• Bedrooms weight barely movesStandardised• Round bowl in the loss• lr 0.1 is stable• Converges in 57 steps• Reaches the noise floor, MSE 22.3
One learning rate cannot suit a steep and a shallow direction at once; scaling makes every direction equally steep.

What you need to know

The stretched valley

Picture the loss as a landscape, with the best weights at the lowest point.

  • If a weight multiplies a feature with huge values (area in sq ft, up to 2,500), a tiny change in that weight changes the loss a lot. That direction is steep.
  • If a weight multiplies a small feature (bedrooms, 1 to 4), its direction is shallow.

One learning rate must serve both. A rate small enough not to overshoot the steep direction crawls along the shallow one. A rate big enough for the shallow direction overshoots the steep one, and the weights blow up.

Measuring it

Here is gradient descent written by hand for a house-price model, run three ways.

Python
import numpy as nprng = np.random.default_rng(0)area = rng.uniform(500, 2500, 500)                  # sq ftbedrooms = rng.integers(1, 5, 500).astype(float)    # 1 to 4price = 20 + 0.06 * area + 3 * bedrooms + rng.normal(0, 5, 500)   # lakhdef gradient_descent(X, y, lr, max_steps=100_000):    X = np.column_stack([np.ones(len(X)), X])       # add intercept column    w = np.zeros(X.shape[1])    for step in range(1, max_steps + 1):        grad = 2 * X.T @ (X @ w - y) / len(y)       # gradient of mean squared error        w -= lr * grad        if not np.isfinite(w).all():            return f"diverged at step {step}"        if np.linalg.norm(grad) < 1e-3:            return f"converged in {step} steps, MSE {np.mean((X @ w - y) ** 2):.1f}"    return f"not converged after {max_steps} steps, MSE {np.mean((X @ w - y) ** 2):.1f}"raw = np.column_stack([area, bedrooms])scaled = (raw - raw.mean(axis=0)) / raw.std(axis=0)with np.errstate(over="ignore", invalid="ignore"):    print("raw,    lr=1e-6:", gradient_descent(raw, price, lr=1e-6))    print("raw,    lr=1e-7:", gradient_descent(raw, price, lr=1e-7))    print("scaled, lr=0.1 :", gradient_descent(scaled, price, lr=0.1))
Text
raw,    lr=1e-6: diverged at step 458raw,    lr=1e-7: not converged after 100000 steps, MSE 122.6scaled, lr=0.1 : converged in 57 steps, MSE 22.3
  • With raw features, a learning rate of one in a million is already too big for the area direction: the weights explode.
  • Ten times smaller is stable but hopeless: after 100,000 steps the error is still 122.6, because the bedrooms and intercept directions barely move.
  • With standardised features, a normal learning rate of 0.1 reaches the best possible error (about 22, the noise we added) in 57 steps.

Numeric safety

Large raw inputs create large activations and gradients in neural networks. That can cause overflow, exploding gradients, or saturated activations such as sigmoid outputs stuck at 0 or 1, where learning stops. Scaled inputs keep numbers in a range where floating-point arithmetic and activations behave well.

A real-life example

A team trains a small neural network for delivery-time prediction with raw features: distance in metres (up to 15,000), order value in rupees (up to 5,000) and a rain flag (0 or 1). The loss becomes nan after two epochs at the default Adam learning rate of 0.001. Lowering the rate stops the nan but training crawls for hours. Adding standardisation to the input pipeline lets training converge in 20 minutes at the default rate. The rain flag, previously ignored, becomes one of the most useful features.

Follow-up questions to expect

  • "Do adaptive optimisers like Adam remove the need for scaling?" — They help, because they adjust the step size per weight, but they do not fully fix badly scaled inputs. Scaling is still standard practice.
  • "What is batch normalisation?" — A layer inside a neural network that standardises activations within each mini-batch during training, which helps deep networks train faster. It does not replace input scaling.
  • "Does scaling change the final answer of linear regression?" — Not the predictions of the unregularised optimum, only how fast gradient descent reaches it and the numeric values of the coefficients.