Course Content
Machine Learning Foundations
14 sections · 70 lessons
What is L1 regularization, and how is it different from L2?
What you need to know
L1 loss = prediction error + λ × (|w₁| + |w₂| + ... + |wₙ|)L2 loss = prediction error + λ × (w₁² + w₂² + ... + wₙ²)The only difference is absolute value versus square — and that small change gives very different behaviour.
Why L1 produces exact zeros
Think about the "push" each penalty gives a weight towards zero:
- L2 pushes harder the larger the weight is, and gently when the weight is small. As a weight shrinks, the push fades, so it rarely reaches zero.
- L1 pushes with the same force at every size. A weight that is not helping enough to beat that constant push goes all the way to zero and stays there.
So L1 acts like a filter: a feature must earn its place, or it is removed.
L1 selecting features
1import numpy as np2from sklearn.datasets import make_regression3from sklearn.linear_model import Ridge, Lasso45# 50 candidate features, only 5 of which really affect the target6X, y, true_coef = make_regression(n_samples=200, n_features=50, n_informative=5,7 noise=10, coef=True, random_state=0)8print("truly useful features:", np.flatnonzero(true_coef))910for name, m in [("Ridge", Ridge(alpha=2.0)), ("Lasso", Lasso(alpha=2.0))]:11 m.fit(X, y)12 kept = np.flatnonzero(m.coef_)13 print(f"{name}: {len(kept)} non-zero coefficients", kept if len(kept) < 10 else "")truly useful features: [ 0 8 23 29 47]Ridge: 50 non-zero coefficients Lasso: 5 non-zero coefficients [ 0 8 23 29 47]Ridge keeps all 50 features with small weights. Lasso keeps exactly the 5 that matter and sets the other 45 to zero. On real data it is rarely this clean, but the behaviour is the same: a smaller, easier-to-explain model.
Side by side
L1 (Lasso)
- Penalty: sum of absolute weights
- Sets many weights to exactly zero
- Built-in feature selection
- From a group of correlated features, keeps one somewhat arbitrarily
L2 (Ridge)
- Penalty: sum of squared weights
- Shrinks all weights, keeps all features
- No selection
- Spreads weight evenly across correlated features
Elastic Net
Elastic Net adds both penalties. In scikit-learn, ElasticNet(alpha=..., l1_ratio=0.5) is half L1, half L2. It keeps L1's ability to drop useless features, but when several features are correlated it tends to keep or drop them together, which is more stable than Lasso's arbitrary pick. For logistic regression in scikit-learn 1.8 and later, you choose the mix with LogisticRegression(l1_ratio=...) — 0 for pure L2, 1 for pure L1 — using a solver such as "saga"; the older penalty="l1" argument is deprecated.
Which to choose
- Many features, most probably useless (hundreds of candidate signals, text n-grams) → L1 or Elastic Net.
- Most features carry some signal, some correlated → L2.
- Need a short list to explain to regulators → L1, then check the kept features make sense.
- Unsure → Elastic Net and let cross-validation pick
l1_ratio.
A real-life example
A bank's credit team starts with 400 candidate features for a loan-default model, many built by different analysts over the years. Regulators require the bank to explain each decision. The team trains a logistic regression with L1 regularisation and tunes the strength with cross-validation; it keeps 28 features and sets the rest to zero, with almost no loss in ROC-AUC. They notice L1 kept "months at current job" but dropped the very similar "months at current employer", so they check both with a domain expert and keep the one that is more reliably reported.
Follow-up questions to expect
- "Can L1 replace proper feature selection?" — It is a good first filter, but with correlated features its choice is unstable. Check the kept features across several cross-validation folds before trusting them.
- "Why is L1 harder to optimise?" — The absolute value has a sharp corner at zero, so the gradient is not defined there. Solvers such as coordinate descent handle it.
- "Does L1 help with multicollinearity?" — Partly — it removes some duplicates — but it picks arbitrarily among them. L2 or Elastic Net handles correlated groups more stably.