Course Content
Machine Learning Foundations
14 sections · 70 lessons
Why does gradient descent converge faster with scaled features?
What you need to know
The stretched valley
Picture the loss as a landscape, with the best weights at the lowest point.
- If a weight multiplies a feature with huge values (area in sq ft, up to 2,500), a tiny change in that weight changes the loss a lot. That direction is steep.
- If a weight multiplies a small feature (bedrooms, 1 to 4), its direction is shallow.
One learning rate must serve both. A rate small enough not to overshoot the steep direction crawls along the shallow one. A rate big enough for the shallow direction overshoots the steep one, and the weights blow up.
Measuring it
Here is gradient descent written by hand for a house-price model, run three ways.
1import numpy as np23rng = np.random.default_rng(0)4area = rng.uniform(500, 2500, 500) # sq ft5bedrooms = rng.integers(1, 5, 500).astype(float) # 1 to 46price = 20 + 0.06 * area + 3 * bedrooms + rng.normal(0, 5, 500) # lakh78def gradient_descent(X, y, lr, max_steps=100_000):9 X = np.column_stack([np.ones(len(X)), X]) # add intercept column10 w = np.zeros(X.shape[1])11 for step in range(1, max_steps + 1):12 grad = 2 * X.T @ (X @ w - y) / len(y) # gradient of mean squared error13 w -= lr * grad14 if not np.isfinite(w).all():15 return f"diverged at step {step}"16 if np.linalg.norm(grad) < 1e-3:17 return f"converged in {step} steps, MSE {np.mean((X @ w - y) ** 2):.1f}"18 return f"not converged after {max_steps} steps, MSE {np.mean((X @ w - y) ** 2):.1f}"1920raw = np.column_stack([area, bedrooms])21scaled = (raw - raw.mean(axis=0)) / raw.std(axis=0)22with np.errstate(over="ignore", invalid="ignore"):23 print("raw, lr=1e-6:", gradient_descent(raw, price, lr=1e-6))24 print("raw, lr=1e-7:", gradient_descent(raw, price, lr=1e-7))25 print("scaled, lr=0.1 :", gradient_descent(scaled, price, lr=0.1))raw, lr=1e-6: diverged at step 458raw, lr=1e-7: not converged after 100000 steps, MSE 122.6scaled, lr=0.1 : converged in 57 steps, MSE 22.3- With raw features, a learning rate of one in a million is already too big for the area direction: the weights explode.
- Ten times smaller is stable but hopeless: after 100,000 steps the error is still 122.6, because the bedrooms and intercept directions barely move.
- With standardised features, a normal learning rate of 0.1 reaches the best possible error (about 22, the noise we added) in 57 steps.
Numeric safety
Large raw inputs create large activations and gradients in neural networks. That can cause overflow, exploding gradients, or saturated activations such as sigmoid outputs stuck at 0 or 1, where learning stops. Scaled inputs keep numbers in a range where floating-point arithmetic and activations behave well.
A real-life example
A team trains a small neural network for delivery-time prediction with raw features: distance in metres (up to 15,000), order value in rupees (up to 5,000) and a rain flag (0 or 1). The loss becomes nan after two epochs at the default Adam learning rate of 0.001. Lowering the rate stops the nan but training crawls for hours. Adding standardisation to the input pipeline lets training converge in 20 minutes at the default rate. The rain flag, previously ignored, becomes one of the most useful features.
Follow-up questions to expect
- "Do adaptive optimisers like Adam remove the need for scaling?" — They help, because they adjust the step size per weight, but they do not fully fix badly scaled inputs. Scaling is still standard practice.
- "What is batch normalisation?" — A layer inside a neural network that standardises activations within each mini-batch during training, which helps deep networks train faster. It does not replace input scaling.
- "Does scaling change the final answer of linear regression?" — Not the predictions of the unregularised optimum, only how fast gradient descent reaches it and the numeric values of the coefficients.