Course Content
Data Science Fundamentals
4 sections · 10 lessons
Encoding Categorical Features
Your model will not train. It says ValueError: could not convert string to float: 'Manchester'. The city column contains text, and every algorithm you have available wants numbers. The obvious fix takes one line:
from sklearn.preprocessing import LabelEncoderdf["city_encoded"] = LabelEncoder().fit_transform(df["city"])Now city is numeric. Birmingham is 0, Bristol is 1, Leeds is 2, London is 3, Manchester is 4. The model trains. It even produces predictions.
Look at what you have told it. A linear model fits one coefficient to this column, so if that coefficient is 1,200 the model has learned that moving from Birmingham to Bristol adds £1,200 to the prediction, and that Manchester adds £4,800 — exactly four times Bristol. It believes Leeds sits precisely halfway between Bristol and London. It believes the average of Birmingham and Leeds is Bristol.
None of this is true. You invented an ordering (alphabetical, as it happens) and invented equal spacing between the categories, and the model has faithfully learned your invention. Nothing errors. The predictions are simply built on nonsense.
Encoding is the problem of turning categories into numbers without smuggling in claims the categories never made.
What Kind of Category Is It?
Everything follows from one question: does the category have a genuine order?
| Type | Example | Is arithmetic meaningful? | Encode with |
|---|---|---|---|
| Nominal — no order | City, colour, payment method | No | One-hot, or target encoding |
| Ordinal — real order | Small < Medium < Large; Bronze < Silver < Gold | Order yes, spacing no | Ordinal encoding with an explicit order |
| Binary | Yes/No, subscribed/not | Trivially | Single 0/1 column |
| High-cardinality nominal | Postcode, product SKU, user ID | No | Target encoding, grouping, or embeddings |
Assigning a number to a category is a claim about order and distance. Make that claim only when it is true.
One-Hot Encoding
Give each category its own column containing 1 if the row belongs to it and 0 otherwise. No ordering is implied because there is no shared axis to be ordered on.
city city_Leeds city_London city_ManchesterLondon -> 0 1 0Manchester -> 0 0 1Leeds -> 1 0 01import pandas as pd23# Quick, for exploration4encoded = pd.get_dummies(df, columns=["city", "plan"], prefix_sep="_", dtype=int)pd.get_dummies is convenient and dangerous for anything beyond exploration, because it derives the columns from whatever values happen to be present. Run it on training data with five cities and on test data with four, and you get frames with different shapes and different column orders. The model then receives features in the wrong positions, which produces wrong answers rather than an error.
1from sklearn.preprocessing import OneHotEncoder23enc = OneHotEncoder(4 handle_unknown="ignore", # unseen category -> all zeros, no crash5 sparse_output=False,6 min_frequency=20, # rare categories folded into one "infrequent" column7)8X_train_enc = enc.fit_transform(X_train[["city", "plan"]])9X_test_enc = enc.transform(X_test[["city", "plan"]]) # same columns, same order10print(enc.get_feature_names_out())handle_unknown="ignore" is what makes this production-safe. A city that never appeared in training becomes a row of zeros — the model treats it as "none of the known cities", which is a reasonable answer — instead of raising an exception in the middle of a live request.
Should you drop one column?
With three cities, the third column is fully determined by the other two: if city_Leeds and city_London are both 0, the row must be Manchester. That redundancy is called the dummy variable trap, and it matters for exactly one family of models.
| Model | Drop a column? | Reason |
|---|---|---|
| Linear/logistic regression with an intercept | Yes — drop="first" | Perfect collinearity makes coefficients unstable or unsolvable |
| Ridge, Lasso | Optional | Regularisation handles the collinearity |
| Trees, random forests, boosting | No | Dropping a column just hides one category from splits |
| Neural networks | No | Not affected |
When you do drop one, the dropped category becomes the baseline and every remaining coefficient is read relative to it. "London adds £3,400 compared with Birmingham" — the comparison is built into the numbers, and forgetting which category was dropped is a reliable way to misread your own model.
The cost
One-hot encoding a column with 40,000 distinct postcodes produces 40,000 columns. Beyond memory, each column is almost entirely zeros, so most carry too few positive examples to estimate anything reliably. As a rough guide: under 15 categories, one-hot without hesitation; 15 to 50, one-hot if you have plenty of rows; above 50, look for something else.
Ordinal Encoding — and How It Differs From Label Encoding
These two get confused constantly, and the confusion is what produced the failure at the top of this lesson.
Ordinal encoding maps categories to integers in an order you specify, for features that genuinely have one:
1from sklearn.preprocessing import OrdinalEncoder23sizes = ["Small", "Medium", "Large", "Extra Large"]4grades = ["Bronze", "Silver", "Gold", "Platinum"]56enc = OrdinalEncoder(7 categories=[sizes, grades],8 handle_unknown="use_encoded_value",9 unknown_value=-1,10)11X_train[["size", "tier"]] = enc.fit_transform(X_train[["size", "tier"]])Passing categories explicitly is the entire point. Left to itself, the encoder sorts alphabetically, so "Extra Large" becomes 0 and "Small" becomes 3 — the order reversed and scrambled, silently.
Even done correctly, integers assert equal spacing. Is the gap from Bronze to Silver the same as from Gold to Platinum? Probably not. Ordinal encoding gets the ordering right and the distances wrong, which is a substantial improvement over getting both wrong, and is usually acceptable.
Label encoding is the same arithmetic with none of the intent. In scikit-learn, LabelEncoder is designed for the target variable — turning class names into integers for a classifier's output — and its documentation says so. Using it on input features is where the alphabetical-ordering bug comes from.
| OrdinalEncoder | LabelEncoder | |
|---|---|---|
| Intended for | Input features with a real order | The target variable |
| Order | You supply it | Alphabetical, always |
| Handles multiple columns | Yes | No — one column at a time |
| Unknown values at predict time | Configurable | Raises an error |
One nuance worth knowing: tree-based models tolerate arbitrary integer codes far better than linear models do, because a tree can carve the numeric axis into many pieces and effectively isolate any category with enough splits. It is inefficient and it costs depth, but it is not catastrophic. For a linear model it is catastrophic, because there is only one coefficient and it must apply to the whole invented ordering.
Target Encoding
Replace each category with a statistic of the target computed within it. Instead of 40,000 postcode columns, one numeric column carrying the signal.
means = df.groupby("postcode")["price"].mean()df["postcode_te"] = df["postcode"].map(means)A postcode where houses average £480,000 becomes 480000; one averaging £190,000 becomes 190000. Dense, informative, and it scales to any cardinality.
It also has two serious failure modes, and the naive version above has both.
Failure one: rare categories
A postcode with a single sale at £2.1 million gets encoded as 2,100,000. That is not the postcode's average price; it is one house. The model sees an extremely high-value feature and learns from noise.
Smoothing fixes it by pulling small groups towards the global average, in proportion to how little evidence they carry:
1def smooth_target_encode(series, target, m=20):2 prior = target.mean()3 stats = target.groupby(series).agg(["mean", "count"])4 smoothed = (stats["count"] * stats["mean"] + m * prior) / (stats["count"] + m)5 return series.map(smoothed).fillna(prior), smoothed, priorWith m = 20 and a global mean of £300,000: a postcode with 1 sale at £2.1M encodes to (1 × 2,100,000 + 20 × 300,000) / 21 = £385,714 — a nudge upwards rather than a wild claim. A postcode with 500 sales averaging £480,000 encodes to £473,077, barely moved, because 500 observations deserve to be believed.
Failure two: leakage
This one is subtler and more damaging. Every row's encoded value was computed using that row's own target. In a category with three members, a row contributes a third of its own feature value. The model can partially read the answer off the input, so training performance looks superb and test performance is poor — and the gap is often blamed on overfitting in the model rather than in the features.
The fix is cross-fitting: encode each fold using only the other folds' targets.
1from sklearn.model_selection import KFold2import numpy as np34def oof_target_encode(series, target, n_splits=5, m=20, seed=42):5 out = np.full(len(series), np.nan)6 kf = KFold(n_splits=n_splits, shuffle=True, random_state=seed)7 for train_idx, val_idx in kf.split(series):8 prior = target.iloc[train_idx].mean()9 stats = target.iloc[train_idx].groupby(series.iloc[train_idx]).agg(["mean", "count"])10 sm = (stats["count"] * stats["mean"] + m * prior) / (stats["count"] + m)11 out[val_idx] = series.iloc[val_idx].map(sm).fillna(prior).to_numpy()12 return out1314X_train["postcode_te"] = oof_target_encode(X_train["postcode"], y_train)The test set is then encoded from the full training data, once. Modern scikit-learn ships TargetEncoder, which performs this cross-fitting internally and is the sensible default:
1from sklearn.preprocessing import TargetEncoder2te = TargetEncoder(smooth="auto", cv=5, random_state=42)3X_train_enc = te.fit_transform(X_train[["postcode"]], y_train)4X_test_enc = te.transform(X_test[["postcode"]])Any encoding that uses the target must be fitted inside the cross-validation loop. Fitted outside it, the validation score measures how well the encoding leaked, not how well the model generalises.
High Cardinality Without the Target
Several approaches avoid target encoding's risks entirely.
1# Frequency encoding: replace with how often the category appears2freq = X_train["sku"].value_counts(normalize=True)3X_train["sku_freq"] = X_train["sku"].map(freq)4X_test["sku_freq"] = X_test["sku"].map(freq).fillna(0)56# Grouping: keep the top N, bucket the rest7top = X_train["sku"].value_counts().nlargest(30).index8X_train["sku_grouped"] = X_train["sku"].where(X_train["sku"].isin(top), "OTHER")910# Domain decomposition: extract meaning from the identifier11df["postcode_area"] = df["postcode"].str.extract(r"^([A-Z]{1,2})") # "SW1A 1AA" -> "SW"That last one is usually the best idea available and the most often overlooked. A postcode, a product code or a timestamp-derived ID typically has structure inside it. Extracting the area, the category prefix or the launch year converts 40,000 categories into 20 meaningful ones without touching the target at all.
| Method | Columns produced | Suits cardinality | Leakage risk | Interpretable |
|---|---|---|---|---|
| One-hot | One per category | Under ~15 | None | High |
| Ordinal (explicit order) | 1 | Any, if ordered | None | High |
| Frequency | 1 | Any | None | Medium |
| Grouping to top-N | N + 1 | Any | None | High |
| Target encoding | 1 | High | High without cross-fitting | Medium |
| Hashing | Fixed (you choose) | Very high | None | None |
Putting It Together
Different columns need different treatment, and the routing belongs in a ColumnTransformer so that every fit/transform boundary is handled for you.
1from sklearn.compose import ColumnTransformer2from sklearn.pipeline import Pipeline3from sklearn.impute import SimpleImputer4from sklearn.preprocessing import OneHotEncoder, OrdinalEncoder, TargetEncoder, StandardScaler5from sklearn.ensemble import HistGradientBoostingRegressor67nominal_low = ["city", "payment_method"]8ordinal_cols = ["size", "tier"]9high_card = ["postcode", "sku"]10numeric = ["age", "income", "quantity"]1112pre = ColumnTransformer([13 ("nom", Pipeline([14 ("impute", SimpleImputer(strategy="constant", fill_value="MISSING")),15 ("ohe", OneHotEncoder(handle_unknown="ignore", min_frequency=20)),16 ]), nominal_low),1718 ("ord", OrdinalEncoder(19 categories=[["Small", "Medium", "Large", "Extra Large"],20 ["Bronze", "Silver", "Gold", "Platinum"]],21 handle_unknown="use_encoded_value", unknown_value=-1,22 ), ordinal_cols),2324 ("high", TargetEncoder(smooth="auto", cv=5, random_state=42), high_card),2526 ("num", Pipeline([27 ("impute", SimpleImputer(strategy="median")),28 ("scale", StandardScaler()),29 ]), numeric),30], remainder="drop")3132model = Pipeline([("pre", pre), ("reg", HistGradientBoostingRegressor(random_state=42))])33model.fit(X_train, y_train)Note how missing categorical values are handled: filled with the literal string "MISSING" rather than dropped or filled with the mode. "We do not know this customer's city" is frequently informative, and one-hot encoding gives it a column of its own to be informative in.
What This Means When You Build Something
Before encoding any column, answer three questions in order. Does it have a real order? If yes, use ordinal encoding with the order written out explicitly. How many distinct values are there? Under fifteen, one-hot and stop thinking about it. Above fifty, look for structure inside the identifier first, then frequency encoding, and only then target encoding with cross-fitting.
Then check what actually came out. The number of features after encoding is the number people are most often surprised by:
names = model.named_steps["pre"].get_feature_names_out()print(f"{X_train.shape[1]} raw columns -> {len(names)} encoded features")print([n for n in names if n.startswith("nom__city")])If twelve raw columns became four thousand features, a high-cardinality column slipped into the one-hot list, and your model is about to spend its capacity on columns that are 99.97% zero.
The consistent theme is that every encoding embeds an assumption. One-hot assumes categories are unrelated. Ordinal assumes an order and equal steps. Target encoding assumes the category's historical average will hold. None of these is always right — but they are all defensible when chosen deliberately, and all indefensible when they arrive by accident because a line of code ran without error.