Course Content
Machine Learning Foundations
14 sections · 70 lessons
How does good feature engineering improve model performance?
What you need to know
A model's job is to find a function from features to the target. Feature engineering improves performance by making that function easier to find. There are five main ways it does this.
- Exposes hidden signal — a raw timestamp says little; "hour of day" and "is holiday" say a lot.
- Simplifies the relationship — a ratio or a log can turn a curved, tangled pattern into a straight line a simple model can learn.
- Reduces noise — averaging over 30 days, capping extreme values and grouping rare categories make the input steadier.
- Adds knowledge the data does not contain — a festival calendar, a list of known risky merchants.
- Needs less data — a model that does not have to discover "peak hour" by itself needs fewer examples to reach the same accuracy.
Relative features beat absolute ones
The most useful trick for beginners is to compare a value with a normal value. A UPI payment of ₹20,000 is ordinary for a business owner and alarming for a student who usually pays ₹200. The raw amount mixes both people together. The ratio amount / user's usual amount puts them on the same scale: 1 means normal, 10 means ten times normal.
1import numpy as np2from sklearn.linear_model import LogisticRegression3from sklearn.model_selection import cross_val_score45rng = np.random.default_rng(1)6n = 200007usual = rng.lognormal(mean=7, sigma=1.5, size=n) # each payer's usual spend, in rupees8is_fraud = rng.random(n) < 0.029ratio = np.where(is_fraud, rng.uniform(2, 12, n), rng.lognormal(0, 0.6, n))10amount = usual * ratio1112X_raw = np.log(amount).reshape(-1, 1) # amount alone13X_eng = np.log(amount / usual).reshape(-1, 1) # amount vs this payer's normal1415for name, X in [("amount", X_raw), ("amount / usual", X_eng)]:16 auc = cross_val_score(LogisticRegression(), X, is_fraud, cv=5, scoring="roc_auc").mean()17 print(f"{name:15s} ROC-AUC = {auc:.2f}")amount ROC-AUC = 0.78amount / usual ROC-AUC = 0.99ROC-AUC measures how well the model ranks fraud above normal payments (0.5 is guessing, 1.0 is perfect; the Threshold Tuning section covers it). Both models see the same information in principle. The second one is handed it in a shape a straight-line model can use. The data here is synthetic and the numbers are cleaner than real life, but the direction of the effect is what you see in practice.
Why this beats algorithm-hopping
Switching from logistic regression to gradient boosting might let the model approximate the ratio on its own, given enough rows. The engineered feature gives it for free, works with a simple model, and can be explained to a compliance team in one sentence: "we flag payments far above the payer's normal". Tuning hyperparameters, by contrast, usually moves a score by a few points at most; it cannot create signal that is not in the features.
A real-life example
A payments company's first fraud model uses amount, time and merchant category. It catches 40% of fraud at a false-alarm rate the operations team can handle. The analysts add three relative features: amount divided by the payer's 90-day median, number of payments in the last 10 minutes, and "is this payee new for this payer". With the same algorithm and the same alert budget, the model now catches 70% of fraud (numbers made up for illustration). The team spends one more week on features and gets more improvement than a month of hyperparameter tuning had given.
Follow-up questions to expect
- "Which models benefit most from feature engineering?" — Simple models like linear and logistic regression benefit most, because they cannot build interactions or curves by themselves. Tree ensembles benefit too, especially from aggregates and ratios.
- "How do you choose which features to build?" — Start from how a domain expert would decide, look at the errors the current model makes, and build features that would explain those errors.
- "How do you avoid building useless features?" — Add them one group at a time and keep only those that improve the validation score by more than the normal run-to-run noise.