Machine Learning Foundations

Course Content

Machine Learning Foundations

14 sections · 70 lessons

What are common feature engineering mistakes in real-world ML projects?


Target encoding a merchant ID that carries no signalNaive fraud rate per merchant• Each row's own labelleaks into its feature• Train AUC 0.95• Test AUC 0.48• Looks brilliant, learned nothingCross-fitted TargetEncoder• Each row encoded from other folds only• Train AUC 0.51• Test AUC 0.52• Honest: the column is noise
A merchant seen once gets its own label as its feature, so the leak scores 0.95 on a column with no signal at all.

What you need to know

Feature mistakes are dangerous because they rarely cause an error message. The code runs, the validation score looks good, and the model fails only in production. Here are the ones interviewers expect you to name, in rough order of damage.

1. Leakage through the feature itself

Leakage means the model sees information during training that it will not have when making real predictions. (The Data Splitting section covers leakage from bad splits.) A classic feature leak: predicting whether a loan will default using num_collection_calls. Collection calls happen after a default, so the feature is almost the answer. Validation looks brilliant; production is useless, because at approval time the value is always zero.

The test: for each feature, ask "at the exact moment of prediction, would I know this value?"

2. Fitting transforms on all the data

Scalers, imputers and encoders learn something (a mean, a category list, a target rate). If they learn it from the whole dataset before splitting, test information leaks into training. The subtle and powerful case is target encoding. Here the category is pure noise, yet the naive version looks excellent on training data:

Python
import numpy as np, pandas as pdfrom sklearn.linear_model import LogisticRegressionfrom sklearn.model_selection import train_test_split, StratifiedKFoldfrom sklearn.metrics import roc_auc_scorefrom sklearn.preprocessing import TargetEncoderrng = np.random.default_rng(0)n = 20000df = pd.DataFrame({"merchant_id": rng.integers(0, 8000, n)})   # 8,000 merchants, ~2 rows eachdf["fraud"] = rng.random(n) < 0.1                               # pure noise: merchant says nothingtr, te = train_test_split(df, random_state=0)# WRONG: fraud rate per merchant, computed from the training labels themselvesrate = tr.groupby("merchant_id").fraud.mean()leaky_tr = tr.merchant_id.map(rate).to_frame()leaky_te = te.merchant_id.map(rate).fillna(tr.fraud.mean()).to_frame()m = LogisticRegression().fit(leaky_tr, tr.fraud)print("naive   train AUC", round(roc_auc_score(tr.fraud, m.predict_proba(leaky_tr)[:, 1]), 2),      "| test AUC", round(roc_auc_score(te.fraud, m.predict_proba(leaky_te)[:, 1]), 2))# RIGHT: TargetEncoder cross-fits, so a row never sees its own labelenc = TargetEncoder(cv=StratifiedKFold(5, shuffle=True, random_state=0))   # 1.9+ stylesafe_tr = enc.fit_transform(tr[["merchant_id"]], tr.fraud)m = LogisticRegression().fit(safe_tr, tr.fraud)print("sklearn train AUC", round(roc_auc_score(tr.fraud, m.predict_proba(safe_tr)[:, 1]), 2),      "| test AUC", round(roc_auc_score(te.fraud, m.predict_proba(enc.transform(te[["merchant_id"]]))[:, 1]), 2))
Text
naive   train AUC 0.95 | test AUC 0.48sklearn train AUC 0.51 | test AUC 0.52

The naive encoding scores 0.95 on training data from a column that contains no signal at all: a merchant with one fraud row gets a rate of 1.0, which is simply its own label. On new data it collapses to 0.48, a coin flip. The cross-fitted version honestly reports about 0.5 on both. The fix in general is to put every learned transform inside a scikit-learn Pipeline, so it is re-fitted on the training part of every split.

3. Training–serving skew

The feature is computed one way in the notebook and another way in the live API. Examples: training uses distance in km, the app sends metres; training fills missing values with the median, the API fills them with 0; training uses end-of-day balances, the API uses the live balance.

4. Ignoring time

Aggregates such as "user's average spend" computed over the full history include the future. Every feature row needs a cutoff time, and splits on time-ordered data should be by time too.

5. The quieter mistakes

  • Blind one-hot on high-cardinality columns — 100,000 sparse columns, slow training, overfitting.
  • Too many features for too few rows — 500 features and 800 rows invites overfitting.
  • Dropping a feature because its correlation with the target is near zero — correlation only sees straight-line relationships; hour of day can have zero correlation with delivery time and still be crucial.
  • No documentation or versioning — six months later nobody knows how risk_score_v2 was computed.

A real-life example

A hospital team builds a model to flag which patients will need intensive care. It scores 0.97 ROC-AUC on validation. A clinician looks at the top features and sees ventilator_ordered — a treatment that is only ordered once a patient is already in intensive care. The model had learned to read the treatment, not predict the need. After the team removes every feature recorded after the prediction time, the score drops to 0.81. That is the honest number, and it is the one that would hold up on the ward.

Follow-up questions to expect

  • "How do you detect leakage?" — Be suspicious of scores that look too good, check which features dominate importance, and for each top feature ask when it is recorded relative to the prediction moment.
  • "How does a Pipeline prevent leakage?" — cross_val_score or GridSearchCV on a Pipeline re-fits the scaler, imputer and encoder on the training folds only, so the validation fold stays unseen.
  • "How do you catch training–serving skew?" — Log the features the live model receives and compare their distributions with the training data; a feature whose average suddenly differs by 1,000 times is usually a unit mismatch.