Course Content
Machine Learning Foundations
14 sections · 70 lessons
How do you handle categorical features in ML models?
What you need to know
A categorical feature takes values from a set of labels rather than a number line: city, payment method, device type. Two properties decide how to encode it:
- Nominal or ordinal? Nominal categories have no order (Pune, Delhi, Chennai). Ordinal categories do (Small, Medium, Large).
- Cardinality — how many distinct values there are. Payment method has 5; merchant ID can have 500,000.
The main encodings
| Encoding | What it does | Use when | Watch out for |
|---|---|---|---|
| One-hot | One 0/1 column per category | Nominal, fewer than ~20–50 values | Thousands of columns for high cardinality |
| Ordinal | S→0, M→1, L→2 | Real order exists | On nominal data it invents a fake order |
| Target (mean) | Replace category with average target for it | High cardinality | Leaks the label unless cross-fitted |
| Frequency | Replace with how often it appears | High cardinality, quick baseline | Two different categories can get the same number |
| Embedding | Learned dense vector per category | Neural nets, IDs with millions of values | Needs lots of data |
One-hot and ordinal in scikit-learn
1import pandas as pd2from sklearn.preprocessing import OneHotEncoder, OrdinalEncoder34train = pd.DataFrame({"city": ["Pune", "Delhi", "Pune", "Chennai"],5 "size": ["S", "L", "M", "S"]})6new = pd.DataFrame({"city": ["Delhi", "Indore"], "size": ["M", "L"]})78ohe = OneHotEncoder(handle_unknown="ignore", sparse_output=False).fit(train[["city"]])9print(ohe.get_feature_names_out())10print(ohe.transform(new[["city"]])) # Indore was never seen in training1112ord_enc = OrdinalEncoder(categories=[["S", "M", "L"]]).fit(train[["size"]])13print(ord_enc.transform(new[["size"]]).ravel())['city_Chennai' 'city_Delhi' 'city_Pune'][[0. 1. 0.] [0. 0. 0.]][1. 2.]Three things to notice. The encoder learns its categories from training data only. handle_unknown="ignore" turns a new city like Indore into all zeros instead of raising an error in production. And for ordinal encoding I passed the order explicitly; left alone, scikit-learn sorts alphabetically (L, M, S), which would be the wrong order.
Why ordinal encoding on cities is wrong
If you map Chennai→0, Delhi→1, Pune→2, a linear model reads that as "Pune is twice Delhi" and "Delhi is between Chennai and Pune". That order is invented. Tree models are less harmed because they can split anywhere, but one-hot is the safe default for nominal data.
High cardinality
For a merchant ID with 200,000 values, one-hot gives 200,000 mostly-zero columns. Better options:
- Group rare values —
OneHotEncoder(min_frequency=50)puts merchants with fewer than 50 rows into one "infrequent" column. - Target encoding — replace each merchant with its fraud rate. scikit-learn's
TargetEncoderdoes this with internal cross-fitting: each training row's value is computed from other folds, so a row never sees its own label. - Native support — LightGBM, CatBoost and scikit-learn's
HistGradientBoostingClassifier(categorical_features="from_dtype")handle categories directly.
A real-life example
A UPI fraud team has payment_app (6 values), device_type (3 values), payer_bank (about 60 values) and merchant_id (400,000 values). They one-hot encode the first two, one-hot payer_bank with rare banks grouped as "other", and target-encode merchant_id with cross-fitting. A new merchant who signs up after training gets the global average fraud rate rather than a crash. When they first tried one-hot on merchant_id, the feature matrix had 400,000 columns and training took hours for no gain.
Follow-up questions to expect
- "What is the dummy variable trap?" — With one-hot plus an intercept in an unregularised linear model, the columns always sum to 1, so they are perfectly collinear. Drop one column (
drop="first") or use regularisation. Tree models do not care. - "How do you handle a category that appears only in the test data?" — Map it to an "unknown" bucket: all zeros for one-hot, the global mean for target encoding.
- "Why does target encoding leak?" — If a category appears once, its "mean target" is exactly that row's label, so the model is handed the answer. Cross-fitting and smoothing toward the global mean prevent this.