Course Content
Machine Learning Foundations
14 sections · 70 lessons
What is the difference between raw features and engineered features?
What you need to know
Think of raw features as ingredients and engineered features as a prepared dish. A model can eat some ingredients directly (a numeric price), but many need preparation first.
| Raw feature | Engineered feature | Why it helps |
|---|---|---|
order_timestamp | hour of day, day of week, is holiday | Behaviour repeats by time, not by exact second |
price | price ÷ category median price | "Expensive" depends on what is being sold |
city (text) | one-hot columns, or average delivery time per city | Models need numbers, not strings |
| event log of clicks | clicks in last 7 days, days since last visit | Turns many rows per user into one row per user |
| review text | sentiment score, text embedding | Turns words into numbers that carry meaning |
From many rows to one row
Most business data is an event log: one row per order, per click, per payment. But a model that predicts something about a user needs one row per user. Aggregation turns the first into the second.
1import pandas as pd23orders = pd.DataFrame({ # raw: one row per order4 "user_id": ["u1", "u1", "u1", "u2", "u2", "u3"],5 "ts": pd.to_datetime(["2026-08-02", "2026-08-20", "2026-09-10",6 "2026-06-15", "2026-09-12", "2026-09-01"]),7 "amount": [450, 1200, 380, 2999, 150, 799],8})9cutoff = pd.Timestamp("2026-09-15") # the moment we predict10past = orders[orders.ts < cutoff] # only data we would have had1112features = past.groupby("user_id").agg( # engineered: one row per user13 orders_total=("amount", "size"),14 avg_amount=("amount", "mean"),15 last_order=("ts", "max"),16)17features["days_since_last"] = (cutoff - features.pop("last_order")).dt.days18features["orders_30d"] = (past[past.ts >= cutoff - pd.Timedelta(days=30)]19 .groupby("user_id").size())20features = features.fillna({"orders_30d": 0})21print(features.round(1)) orders_total avg_amount days_since_last orders_30duser_id u1 3 676.7 5 2u2 2 1574.5 3 1u3 1 799.0 14 1Notice the cutoff line. Every feature is computed only from orders before the moment of prediction. That single line is what keeps the feature honest; without it, a training row could include orders that happened after the event you are predicting.
Where deep learning changes the picture
For images, audio and text, deep networks learn their own features from raw pixels and tokens, so hand-built features matter much less. A CNN learns edge detectors; a transformer learns word meanings. For tabular business data — payments, orders, bookings — hand-built features remain the main lever, and gradient-boosted trees on good features are still the usual strong baseline.
The price of engineered features
Every engineered feature is code that must run twice: once in the training pipeline and once in the live service. If "orders in the last 30 days" is computed from a nightly batch in training but from a live database in production, the two numbers will differ. This is called training–serving skew, and it is one of the most common reasons a model works offline and fails live.
A real-life example
An e-commerce company wants to predict which customers will stop buying (churn). The raw data is 40 million order rows. The first attempt joins the latest order to each customer and scores poorly. The second attempt builds per-customer features: days since last order, orders in the last 30 and 90 days, average basket value, share of orders that were returned, and whether the last delivery was late. "Days since last order" and "late last delivery" turn out to be the two strongest signals — neither exists as a raw column. The team puts the feature code in one shared library so the nightly training job and the live scoring API compute exactly the same values.
Follow-up questions to expect
- "Is one-hot encoding feature engineering?" — Yes, it is a transformation of a raw feature. The more valuable kind of engineering, though, adds new information, like a ratio or a time window.
- "How do you keep training and serving features consistent?" — Use one shared feature code path or a feature store, and log the features the live model actually received so you can compare them with training.
- "What is a point-in-time correct feature?" — A feature computed using only data that existed at the moment of prediction, as the
cutoffline in the code does.