Machine Learning Foundations

Course Content

Machine Learning Foundations

14 sections · 70 lessons

How do you handle categorical features in ML models?


One-hot encoding, including a city never seen in training001010100000city_Chennaicity_Delhicity_PunePuneDelhiChennaiIndore (new)handle_unknown='ignore' turns the new city into all zeros instead of an error.
One column per known category means no invented order between cities, and a safe all-zero row for anything new.

What you need to know

A categorical feature takes values from a set of labels rather than a number line: city, payment method, device type. Two properties decide how to encode it:

  • Nominal or ordinal? Nominal categories have no order (Pune, Delhi, Chennai). Ordinal categories do (Small, Medium, Large).
  • Cardinality — how many distinct values there are. Payment method has 5; merchant ID can have 500,000.

The main encodings

EncodingWhat it doesUse whenWatch out for
One-hotOne 0/1 column per categoryNominal, fewer than ~20–50 valuesThousands of columns for high cardinality
OrdinalS→0, M→1, L→2Real order existsOn nominal data it invents a fake order
Target (mean)Replace category with average target for itHigh cardinalityLeaks the label unless cross-fitted
FrequencyReplace with how often it appearsHigh cardinality, quick baselineTwo different categories can get the same number
EmbeddingLearned dense vector per categoryNeural nets, IDs with millions of valuesNeeds lots of data

One-hot and ordinal in scikit-learn

Python
import pandas as pdfrom sklearn.preprocessing import OneHotEncoder, OrdinalEncodertrain = pd.DataFrame({"city": ["Pune", "Delhi", "Pune", "Chennai"],                      "size": ["S", "L", "M", "S"]})new = pd.DataFrame({"city": ["Delhi", "Indore"], "size": ["M", "L"]})ohe = OneHotEncoder(handle_unknown="ignore", sparse_output=False).fit(train[["city"]])print(ohe.get_feature_names_out())print(ohe.transform(new[["city"]]))          # Indore was never seen in trainingord_enc = OrdinalEncoder(categories=[["S", "M", "L"]]).fit(train[["size"]])print(ord_enc.transform(new[["size"]]).ravel())
Text
['city_Chennai' 'city_Delhi' 'city_Pune'][[0. 1. 0.] [0. 0. 0.]][1. 2.]

Three things to notice. The encoder learns its categories from training data only. handle_unknown="ignore" turns a new city like Indore into all zeros instead of raising an error in production. And for ordinal encoding I passed the order explicitly; left alone, scikit-learn sorts alphabetically (L, M, S), which would be the wrong order.

Why ordinal encoding on cities is wrong

If you map Chennai→0, Delhi→1, Pune→2, a linear model reads that as "Pune is twice Delhi" and "Delhi is between Chennai and Pune". That order is invented. Tree models are less harmed because they can split anywhere, but one-hot is the safe default for nominal data.

High cardinality

For a merchant ID with 200,000 values, one-hot gives 200,000 mostly-zero columns. Better options:

  • Group rare values — OneHotEncoder(min_frequency=50) puts merchants with fewer than 50 rows into one "infrequent" column.
  • Target encoding — replace each merchant with its fraud rate. scikit-learn's TargetEncoder does this with internal cross-fitting: each training row's value is computed from other folds, so a row never sees its own label.
  • Native support — LightGBM, CatBoost and scikit-learn's HistGradientBoostingClassifier(categorical_features="from_dtype") handle categories directly.

A real-life example

A UPI fraud team has payment_app (6 values), device_type (3 values), payer_bank (about 60 values) and merchant_id (400,000 values). They one-hot encode the first two, one-hot payer_bank with rare banks grouped as "other", and target-encode merchant_id with cross-fitting. A new merchant who signs up after training gets the global average fraud rate rather than a crash. When they first tried one-hot on merchant_id, the feature matrix had 400,000 columns and training took hours for no gain.

Follow-up questions to expect

  • "What is the dummy variable trap?" — With one-hot plus an intercept in an unregularised linear model, the columns always sum to 1, so they are perfectly collinear. Drop one column (drop="first") or use regularisation. Tree models do not care.
  • "How do you handle a category that appears only in the test data?" — Map it to an "unknown" bucket: all zeros for one-hot, the global mean for target encoding.
  • "Why does target encoding leak?" — If a category appears once, its "mean target" is exactly that row's label, so the model is handed the answer. Cross-fitting and smoothing toward the global mean prevent this.