Course Content
Machine Learning Foundations
14 sections · 70 lessons
What are common feature engineering mistakes in real-world ML projects?
What you need to know
Feature mistakes are dangerous because they rarely cause an error message. The code runs, the validation score looks good, and the model fails only in production. Here are the ones interviewers expect you to name, in rough order of damage.
1. Leakage through the feature itself
Leakage means the model sees information during training that it will not have when making real predictions. (The Data Splitting section covers leakage from bad splits.) A classic feature leak: predicting whether a loan will default using num_collection_calls. Collection calls happen after a default, so the feature is almost the answer. Validation looks brilliant; production is useless, because at approval time the value is always zero.
The test: for each feature, ask "at the exact moment of prediction, would I know this value?"
2. Fitting transforms on all the data
Scalers, imputers and encoders learn something (a mean, a category list, a target rate). If they learn it from the whole dataset before splitting, test information leaks into training. The subtle and powerful case is target encoding. Here the category is pure noise, yet the naive version looks excellent on training data:
1import numpy as np, pandas as pd2from sklearn.linear_model import LogisticRegression3from sklearn.model_selection import train_test_split, StratifiedKFold4from sklearn.metrics import roc_auc_score5from sklearn.preprocessing import TargetEncoder67rng = np.random.default_rng(0)8n = 200009df = pd.DataFrame({"merchant_id": rng.integers(0, 8000, n)}) # 8,000 merchants, ~2 rows each10df["fraud"] = rng.random(n) < 0.1 # pure noise: merchant says nothing11tr, te = train_test_split(df, random_state=0)1213# WRONG: fraud rate per merchant, computed from the training labels themselves14rate = tr.groupby("merchant_id").fraud.mean()15leaky_tr = tr.merchant_id.map(rate).to_frame()16leaky_te = te.merchant_id.map(rate).fillna(tr.fraud.mean()).to_frame()17m = LogisticRegression().fit(leaky_tr, tr.fraud)18print("naive train AUC", round(roc_auc_score(tr.fraud, m.predict_proba(leaky_tr)[:, 1]), 2),19 "| test AUC", round(roc_auc_score(te.fraud, m.predict_proba(leaky_te)[:, 1]), 2))2021# RIGHT: TargetEncoder cross-fits, so a row never sees its own label22enc = TargetEncoder(cv=StratifiedKFold(5, shuffle=True, random_state=0)) # 1.9+ style23safe_tr = enc.fit_transform(tr[["merchant_id"]], tr.fraud)24m = LogisticRegression().fit(safe_tr, tr.fraud)25print("sklearn train AUC", round(roc_auc_score(tr.fraud, m.predict_proba(safe_tr)[:, 1]), 2),26 "| test AUC", round(roc_auc_score(te.fraud, m.predict_proba(enc.transform(te[["merchant_id"]]))[:, 1]), 2))naive train AUC 0.95 | test AUC 0.48sklearn train AUC 0.51 | test AUC 0.52The naive encoding scores 0.95 on training data from a column that contains no signal at all: a merchant with one fraud row gets a rate of 1.0, which is simply its own label. On new data it collapses to 0.48, a coin flip. The cross-fitted version honestly reports about 0.5 on both. The fix in general is to put every learned transform inside a scikit-learn Pipeline, so it is re-fitted on the training part of every split.
3. Training–serving skew
The feature is computed one way in the notebook and another way in the live API. Examples: training uses distance in km, the app sends metres; training fills missing values with the median, the API fills them with 0; training uses end-of-day balances, the API uses the live balance.
4. Ignoring time
Aggregates such as "user's average spend" computed over the full history include the future. Every feature row needs a cutoff time, and splits on time-ordered data should be by time too.
5. The quieter mistakes
- Blind one-hot on high-cardinality columns — 100,000 sparse columns, slow training, overfitting.
- Too many features for too few rows — 500 features and 800 rows invites overfitting.
- Dropping a feature because its correlation with the target is near zero — correlation only sees straight-line relationships; hour of day can have zero correlation with delivery time and still be crucial.
- No documentation or versioning — six months later nobody knows how
risk_score_v2was computed.
A real-life example
A hospital team builds a model to flag which patients will need intensive care. It scores 0.97 ROC-AUC on validation. A clinician looks at the top features and sees ventilator_ordered — a treatment that is only ordered once a patient is already in intensive care. The model had learned to read the treatment, not predict the need. After the team removes every feature recorded after the prediction time, the score drops to 0.81. That is the honest number, and it is the one that would hold up on the ward.
Follow-up questions to expect
- "How do you detect leakage?" — Be suspicious of scores that look too good, check which features dominate importance, and for each top feature ask when it is recorded relative to the prediction moment.
- "How does a Pipeline prevent leakage?" —
cross_val_scoreorGridSearchCVon aPipelinere-fits the scaler, imputer and encoder on the training folds only, so the validation fold stays unseen. - "How do you catch training–serving skew?" — Log the features the live model receives and compare their distributions with the training data; a feature whose average suddenly differs by 1,000 times is usually a unit mismatch.