Data Science Fundamentals

Encoding Categorical Features


Your model will not train. It says ValueError: could not convert string to float: 'Manchester'. The city column contains text, and every algorithm you have available wants numbers. The obvious fix takes one line:

Python
from sklearn.preprocessing import LabelEncoderdf["city_encoded"] = LabelEncoder().fit_transform(df["city"])

Now city is numeric. Birmingham is 0, Bristol is 1, Leeds is 2, London is 3, Manchester is 4. The model trains. It even produces predictions.

Look at what you have told it. A linear model fits one coefficient to this column, so if that coefficient is 1,200 the model has learned that moving from Birmingham to Bristol adds £1,200 to the prediction, and that Manchester adds £4,800 — exactly four times Bristol. It believes Leeds sits precisely halfway between Bristol and London. It believes the average of Birmingham and Leeds is Bristol.

None of this is true. You invented an ordering (alphabetical, as it happens) and invented equal spacing between the categories, and the model has faithfully learned your invention. Nothing errors. The predictions are simply built on nonsense.

Encoding is the problem of turning categories into numbers without smuggling in claims the categories never made.

One-hot or target encoding, and what each costsOne-hot• No fake ordering is invented• Leaks nothing from the target• One column per level,so 3,000 cities explode• Trees struggle with the sparse widthTarget encoding• One column no matter the cardinality• Carries real signal about the outcome• Rare levels memorise a single row• Leaks unless fitted inside the folds
Label-encoding the city column takes one line and tells the model that Manchester is greater than Leeds.

What Kind of Category Is It?

Everything follows from one question: does the category have a genuine order?

TypeExampleIs arithmetic meaningful?Encode with
Nominal — no orderCity, colour, payment methodNoOne-hot, or target encoding
Ordinal — real orderSmall < Medium < Large; Bronze < Silver < GoldOrder yes, spacing noOrdinal encoding with an explicit order
BinaryYes/No, subscribed/notTriviallySingle 0/1 column
High-cardinality nominalPostcode, product SKU, user IDNoTarget encoding, grouping, or embeddings

Assigning a number to a category is a claim about order and distance. Make that claim only when it is true.

One-Hot Encoding

Give each category its own column containing 1 if the row belongs to it and 0 otherwise. No ordering is implied because there is no shared axis to be ordered on.

Text
city          city_Leeds  city_London  city_ManchesterLondon    ->      0            1              0Manchester ->     0            0              1Leeds     ->      1            0              0
Python
import pandas as pd# Quick, for explorationencoded = pd.get_dummies(df, columns=["city", "plan"], prefix_sep="_", dtype=int)

pd.get_dummies is convenient and dangerous for anything beyond exploration, because it derives the columns from whatever values happen to be present. Run it on training data with five cities and on test data with four, and you get frames with different shapes and different column orders. The model then receives features in the wrong positions, which produces wrong answers rather than an error.

Python
from sklearn.preprocessing import OneHotEncoderenc = OneHotEncoder(    handle_unknown="ignore",   # unseen category -> all zeros, no crash    sparse_output=False,    min_frequency=20,          # rare categories folded into one "infrequent" column)X_train_enc = enc.fit_transform(X_train[["city", "plan"]])X_test_enc = enc.transform(X_test[["city", "plan"]])   # same columns, same orderprint(enc.get_feature_names_out())

handle_unknown="ignore" is what makes this production-safe. A city that never appeared in training becomes a row of zeros — the model treats it as "none of the known cities", which is a reasonable answer — instead of raising an exception in the middle of a live request.

Should you drop one column?

With three cities, the third column is fully determined by the other two: if city_Leeds and city_London are both 0, the row must be Manchester. That redundancy is called the dummy variable trap, and it matters for exactly one family of models.

ModelDrop a column?Reason
Linear/logistic regression with an interceptYes — drop="first"Perfect collinearity makes coefficients unstable or unsolvable
Ridge, LassoOptionalRegularisation handles the collinearity
Trees, random forests, boostingNoDropping a column just hides one category from splits
Neural networksNoNot affected

When you do drop one, the dropped category becomes the baseline and every remaining coefficient is read relative to it. "London adds £3,400 compared with Birmingham" — the comparison is built into the numbers, and forgetting which category was dropped is a reliable way to misread your own model.

The cost

One-hot encoding a column with 40,000 distinct postcodes produces 40,000 columns. Beyond memory, each column is almost entirely zeros, so most carry too few positive examples to estimate anything reliably. As a rough guide: under 15 categories, one-hot without hesitation; 15 to 50, one-hot if you have plenty of rows; above 50, look for something else.

Ordinal Encoding — and How It Differs From Label Encoding

These two get confused constantly, and the confusion is what produced the failure at the top of this lesson.

Ordinal encoding maps categories to integers in an order you specify, for features that genuinely have one:

Python
from sklearn.preprocessing import OrdinalEncodersizes = ["Small", "Medium", "Large", "Extra Large"]grades = ["Bronze", "Silver", "Gold", "Platinum"]enc = OrdinalEncoder(    categories=[sizes, grades],    handle_unknown="use_encoded_value",    unknown_value=-1,)X_train[["size", "tier"]] = enc.fit_transform(X_train[["size", "tier"]])

Passing categories explicitly is the entire point. Left to itself, the encoder sorts alphabetically, so "Extra Large" becomes 0 and "Small" becomes 3 — the order reversed and scrambled, silently.

Even done correctly, integers assert equal spacing. Is the gap from Bronze to Silver the same as from Gold to Platinum? Probably not. Ordinal encoding gets the ordering right and the distances wrong, which is a substantial improvement over getting both wrong, and is usually acceptable.

Label encoding is the same arithmetic with none of the intent. In scikit-learn, LabelEncoder is designed for the target variable — turning class names into integers for a classifier's output — and its documentation says so. Using it on input features is where the alphabetical-ordering bug comes from.

OrdinalEncoderLabelEncoder
Intended forInput features with a real orderThe target variable
OrderYou supply itAlphabetical, always
Handles multiple columnsYesNo — one column at a time
Unknown values at predict timeConfigurableRaises an error

One nuance worth knowing: tree-based models tolerate arbitrary integer codes far better than linear models do, because a tree can carve the numeric axis into many pieces and effectively isolate any category with enough splits. It is inefficient and it costs depth, but it is not catastrophic. For a linear model it is catastrophic, because there is only one coefficient and it must apply to the whole invented ordering.

Target Encoding

Replace each category with a statistic of the target computed within it. Instead of 40,000 postcode columns, one numeric column carrying the signal.

Python
means = df.groupby("postcode")["price"].mean()df["postcode_te"] = df["postcode"].map(means)

A postcode where houses average £480,000 becomes 480000; one averaging £190,000 becomes 190000. Dense, informative, and it scales to any cardinality.

It also has two serious failure modes, and the naive version above has both.

Failure one: rare categories

A postcode with a single sale at £2.1 million gets encoded as 2,100,000. That is not the postcode's average price; it is one house. The model sees an extremely high-value feature and learns from noise.

Smoothing fixes it by pulling small groups towards the global average, in proportion to how little evidence they carry:

y^c=nc⋅yˉc+m⋅yˉnc+m\hat{y}_c = \frac{n_c \cdot \bar{y}_c + m \cdot \bar{y}}{n_c + m}

Python
def smooth_target_encode(series, target, m=20):    prior = target.mean()    stats = target.groupby(series).agg(["mean", "count"])    smoothed = (stats["count"] * stats["mean"] + m * prior) / (stats["count"] + m)    return series.map(smoothed).fillna(prior), smoothed, prior

With m = 20 and a global mean of £300,000: a postcode with 1 sale at £2.1M encodes to (1 × 2,100,000 + 20 × 300,000) / 21 = £385,714 — a nudge upwards rather than a wild claim. A postcode with 500 sales averaging £480,000 encodes to £473,077, barely moved, because 500 observations deserve to be believed.

Failure two: leakage

This one is subtler and more damaging. Every row's encoded value was computed using that row's own target. In a category with three members, a row contributes a third of its own feature value. The model can partially read the answer off the input, so training performance looks superb and test performance is poor — and the gap is often blamed on overfitting in the model rather than in the features.

The fix is cross-fitting: encode each fold using only the other folds' targets.

Python
from sklearn.model_selection import KFoldimport numpy as npdef oof_target_encode(series, target, n_splits=5, m=20, seed=42):    out = np.full(len(series), np.nan)    kf = KFold(n_splits=n_splits, shuffle=True, random_state=seed)    for train_idx, val_idx in kf.split(series):        prior = target.iloc[train_idx].mean()        stats = target.iloc[train_idx].groupby(series.iloc[train_idx]).agg(["mean", "count"])        sm = (stats["count"] * stats["mean"] + m * prior) / (stats["count"] + m)        out[val_idx] = series.iloc[val_idx].map(sm).fillna(prior).to_numpy()    return outX_train["postcode_te"] = oof_target_encode(X_train["postcode"], y_train)

The test set is then encoded from the full training data, once. Modern scikit-learn ships TargetEncoder, which performs this cross-fitting internally and is the sensible default:

Python
from sklearn.preprocessing import TargetEncoderte = TargetEncoder(smooth="auto", cv=5, random_state=42)X_train_enc = te.fit_transform(X_train[["postcode"]], y_train)X_test_enc = te.transform(X_test[["postcode"]])

Any encoding that uses the target must be fitted inside the cross-validation loop. Fitted outside it, the validation score measures how well the encoding leaked, not how well the model generalises.

High Cardinality Without the Target

Several approaches avoid target encoding's risks entirely.

Python
# Frequency encoding: replace with how often the category appearsfreq = X_train["sku"].value_counts(normalize=True)X_train["sku_freq"] = X_train["sku"].map(freq)X_test["sku_freq"] = X_test["sku"].map(freq).fillna(0)# Grouping: keep the top N, bucket the resttop = X_train["sku"].value_counts().nlargest(30).indexX_train["sku_grouped"] = X_train["sku"].where(X_train["sku"].isin(top), "OTHER")# Domain decomposition: extract meaning from the identifierdf["postcode_area"] = df["postcode"].str.extract(r"^([A-Z]{1,2})")   # "SW1A 1AA" -> "SW"

That last one is usually the best idea available and the most often overlooked. A postcode, a product code or a timestamp-derived ID typically has structure inside it. Extracting the area, the category prefix or the launch year converts 40,000 categories into 20 meaningful ones without touching the target at all.

MethodColumns producedSuits cardinalityLeakage riskInterpretable
One-hotOne per categoryUnder ~15NoneHigh
Ordinal (explicit order)1Any, if orderedNoneHigh
Frequency1AnyNoneMedium
Grouping to top-NN + 1AnyNoneHigh
Target encoding1HighHigh without cross-fittingMedium
HashingFixed (you choose)Very highNoneNone

Putting It Together

Different columns need different treatment, and the routing belongs in a ColumnTransformer so that every fit/transform boundary is handled for you.

Python
from sklearn.compose import ColumnTransformerfrom sklearn.pipeline import Pipelinefrom sklearn.impute import SimpleImputerfrom sklearn.preprocessing import OneHotEncoder, OrdinalEncoder, TargetEncoder, StandardScalerfrom sklearn.ensemble import HistGradientBoostingRegressornominal_low = ["city", "payment_method"]ordinal_cols = ["size", "tier"]high_card = ["postcode", "sku"]numeric = ["age", "income", "quantity"]pre = ColumnTransformer([    ("nom", Pipeline([        ("impute", SimpleImputer(strategy="constant", fill_value="MISSING")),        ("ohe", OneHotEncoder(handle_unknown="ignore", min_frequency=20)),    ]), nominal_low),    ("ord", OrdinalEncoder(        categories=[["Small", "Medium", "Large", "Extra Large"],                    ["Bronze", "Silver", "Gold", "Platinum"]],        handle_unknown="use_encoded_value", unknown_value=-1,    ), ordinal_cols),    ("high", TargetEncoder(smooth="auto", cv=5, random_state=42), high_card),    ("num", Pipeline([        ("impute", SimpleImputer(strategy="median")),        ("scale", StandardScaler()),    ]), numeric),], remainder="drop")model = Pipeline([("pre", pre), ("reg", HistGradientBoostingRegressor(random_state=42))])model.fit(X_train, y_train)

Note how missing categorical values are handled: filled with the literal string "MISSING" rather than dropped or filled with the mode. "We do not know this customer's city" is frequently informative, and one-hot encoding gives it a column of its own to be informative in.

What This Means When You Build Something

Before encoding any column, answer three questions in order. Does it have a real order? If yes, use ordinal encoding with the order written out explicitly. How many distinct values are there? Under fifteen, one-hot and stop thinking about it. Above fifty, look for structure inside the identifier first, then frequency encoding, and only then target encoding with cross-fitting.

Then check what actually came out. The number of features after encoding is the number people are most often surprised by:

Python
names = model.named_steps["pre"].get_feature_names_out()print(f"{X_train.shape[1]} raw columns -> {len(names)} encoded features")print([n for n in names if n.startswith("nom__city")])

If twelve raw columns became four thousand features, a high-cardinality column slipped into the one-hot list, and your model is about to spend its capacity on columns that are 99.97% zero.

The consistent theme is that every encoding embeds an assumption. One-hot assumes categories are unrelated. Ordinal assumes an order and equal steps. Target encoding assumes the category's historical average will hold. None of these is always right — but they are all defensible when chosen deliberately, and all indefensible when they arrive by accident because a line of code ran without error.