Course Content
Machine Learning Foundations
14 sections · 70 lessons
What is feature engineering, and why does it matter?
What you need to know
A feature is one input column the model sees. A label (or target) is the answer it must predict. Raw data rarely arrives in a form a model can use well. A timestamp, a free-text address or a log of 200 clicks is information, but it is not yet a useful signal.
Why the model cannot do it alone
A model learns a certain shape of relationship. Linear regression learns straight lines: "one more unit of X adds a fixed amount to Y". If the real pattern is "lunch and dinner hours are slow", a raw hour column from 0 to 23 cannot express that as a straight line — hour 13 is slow, hour 16 is fast, hour 20 is slow again. Give the model an is_peak column (1 at lunch and dinner, 0 otherwise) and the pattern becomes a single number it can learn.
Trees and neural networks are more flexible and can find some of these shapes themselves, but they need more data to do so, and they still cannot use information you never gave them.
A small experiment
Here is a delivery-time model trained twice on the same synthetic orders: once with the raw hour, once with an engineered is_peak flag.
1import numpy as np2from sklearn.linear_model import LinearRegression3from sklearn.model_selection import train_test_split4from sklearn.metrics import mean_absolute_error56rng = np.random.default_rng(0)7n = 50008hour = rng.integers(0, 24, n) # hour the order was placed9distance = rng.uniform(0.5, 8, n) # km from restaurant10is_peak = np.isin(hour, [12, 13, 19, 20, 21]).astype(int)11minutes = 15 + 3 * distance + 12 * is_peak + rng.normal(0, 3, n)1213X_raw = np.column_stack([hour, distance]) # raw columns14X_eng = np.column_stack([is_peak, distance]) # one engineered column1516for name, X in [("raw hour", X_raw), ("is_peak", X_eng)]:17 X_tr, X_te, y_tr, y_te = train_test_split(X, minutes, random_state=0)18 model = LinearRegression().fit(X_tr, y_tr)19 print(f"{name:9s} MAE = {mean_absolute_error(y_te, model.predict(X_te)):.1f} min")raw hour MAE = 4.4 minis_peak MAE = 2.4 minSame algorithm, same rows. The only change is one column, and the average error (MAE, mean absolute error) drops from 4.4 to 2.4 minutes. The remaining 2.4 minutes is the random noise we put in, so the engineered model has found almost all the real signal.
Where good features come from
- Domain knowledge — a delivery partner knows rain and peak hours matter; a fraud analyst knows "first payment to a new payee" matters.
- Time — hour, day of week, holiday flag, days since last event.
- Aggregates — counts, sums and averages over a window ("payments in the last 10 minutes").
- Ratios and differences — amount divided by the user's usual amount, price minus category average.
A real-life example
A food-delivery app in Bengaluru predicts delivery time from order_timestamp, restaurant_id and customer_location. The first model, trained on these raw columns, is off by 9 minutes on average, and customers complain the ETA is always wrong on Friday nights.
The engineer adds five features: hour of day, is weekend, distance in km (computed from the two locations), the restaurant's average preparation time over the last 7 days, and the number of open orders in that area right now. The error drops to 5 minutes without changing the algorithm. The biggest win is "open orders in the area", because it captures the Friday-night rush the raw timestamp could only hint at.
Follow-up questions to expect
- "Does deep learning remove the need for feature engineering?" — Largely, for images, audio and text, where networks learn features from raw pixels and tokens. For tabular data such as payments or orders, hand-built features are still the main lever.
- "How do you know a feature helped?" — Compare the validation score with and without it, using the same split. Also check feature importance, and make sure the feature is available at prediction time.
- "Can you have too many features?" — Yes. With few rows, many weak features lead to overfitting and slower training. Remove the ones that add nothing on validation.