Machine Learning Foundations

Course Content

Machine Learning Foundations

14 sections · 70 lessons

What is the difference between raw features and engineered features?


What you need to know

Think of raw features as ingredients and engineered features as a prepared dish. A model can eat some ingredients directly (a numeric price), but many need preparation first.

Raw featureEngineered featureWhy it helps
order_timestamphour of day, day of week, is holidayBehaviour repeats by time, not by exact second
priceprice ÷ category median price"Expensive" depends on what is being sold
city (text)one-hot columns, or average delivery time per cityModels need numbers, not strings
event log of clicksclicks in last 7 days, days since last visitTurns many rows per user into one row per user
review textsentiment score, text embeddingTurns words into numbers that carry meaning

From many rows to one row

Most business data is an event log: one row per order, per click, per payment. But a model that predicts something about a user needs one row per user. Aggregation turns the first into the second.

Python
import pandas as pdorders = pd.DataFrame({                      # raw: one row per order    "user_id": ["u1", "u1", "u1", "u2", "u2", "u3"],    "ts": pd.to_datetime(["2026-08-02", "2026-08-20", "2026-09-10",                          "2026-06-15", "2026-09-12", "2026-09-01"]),    "amount": [450, 1200, 380, 2999, 150, 799],})cutoff = pd.Timestamp("2026-09-15")          # the moment we predictpast = orders[orders.ts < cutoff]            # only data we would have hadfeatures = past.groupby("user_id").agg(      # engineered: one row per user    orders_total=("amount", "size"),    avg_amount=("amount", "mean"),    last_order=("ts", "max"),)features["days_since_last"] = (cutoff - features.pop("last_order")).dt.daysfeatures["orders_30d"] = (past[past.ts >= cutoff - pd.Timedelta(days=30)]                          .groupby("user_id").size())features = features.fillna({"orders_30d": 0})print(features.round(1))
Text
         orders_total  avg_amount  days_since_last  orders_30duser_id                                                       u1                  3       676.7                5           2u2                  2      1574.5                3           1u3                  1       799.0               14           1

Notice the cutoff line. Every feature is computed only from orders before the moment of prediction. That single line is what keeps the feature honest; without it, a training row could include orders that happened after the event you are predicting.

Where deep learning changes the picture

For images, audio and text, deep networks learn their own features from raw pixels and tokens, so hand-built features matter much less. A CNN learns edge detectors; a transformer learns word meanings. For tabular business data — payments, orders, bookings — hand-built features remain the main lever, and gradient-boosted trees on good features are still the usual strong baseline.

The price of engineered features

Every engineered feature is code that must run twice: once in the training pipeline and once in the live service. If "orders in the last 30 days" is computed from a nightly batch in training but from a live database in production, the two numbers will differ. This is called training–serving skew, and it is one of the most common reasons a model works offline and fails live.

A real-life example

An e-commerce company wants to predict which customers will stop buying (churn). The raw data is 40 million order rows. The first attempt joins the latest order to each customer and scores poorly. The second attempt builds per-customer features: days since last order, orders in the last 30 and 90 days, average basket value, share of orders that were returned, and whether the last delivery was late. "Days since last order" and "late last delivery" turn out to be the two strongest signals — neither exists as a raw column. The team puts the feature code in one shared library so the nightly training job and the live scoring API compute exactly the same values.

Follow-up questions to expect

  • "Is one-hot encoding feature engineering?" — Yes, it is a transformation of a raw feature. The more valuable kind of engineering, though, adds new information, like a ratio or a time window.
  • "How do you keep training and serving features consistent?" — Use one shared feature code path or a feature store, and log the features the live model actually received so you can compare them with training.
  • "What is a point-in-time correct feature?" — A feature computed using only data that existed at the moment of prediction, as the cutoff line in the code does.