Course Content
Statistics & Math for AI/ML Interviews
8 sections · 30 lessons
What is the mode, and where does it appear in classification problems?
What you need to know
The idea
Count how often each value appears and pick the most common one. For the labels ham, ham, spam, ham, spam, ham, the counts are ham = 4 and spam = 2, so the mode is ham.
A dataset can have two modes (bimodal) or more. It can also have no useful mode when every value is different, which is common for continuous numbers like 31.42 and 31.47. For continuous data, the mode is read from the peak of a histogram instead.
Where the mode appears in classification
- Baseline and imbalance check. The mode's share of the labels is your baseline accuracy. If it is 95%, accuracy is the wrong metric; use precision, recall, F1 or PR-AUC.
- Imputing categorical features. A missing "city" can be filled with the most common city, though an explicit "Unknown" category is often better because missingness can carry signal.
- Hard voting. In a voting ensemble, each model votes for a class and the prediction is the mode of the votes. Breiman's original random forest works this way. Note that scikit-learn's
RandomForestClassifieractually averages the trees' predicted probabilities (soft voting) and picks the highest. - k-nearest neighbours. A k-NN classifier predicts the mode of the labels of the k closest points.
1import pandas as pd23labels = pd.Series(["ham"] * 950 + ["spam"] * 50)4print(labels.value_counts())5print("mode:", labels.mode()[0])6print("baseline accuracy:", labels.value_counts(normalize=True).max())ham 950spam 50Name: count, dtype: int64mode: hambaseline accuracy: 0.95value_counts() is the first line most people run on a new label column. mode() returns a Series because there can be ties, which is why the code takes [0].
A real-life example
A bank trains a fraud model on UPI transactions. 99.8% of the labels are "not fraud", so the mode's share is 0.998. The first model reports 99.8% accuracy, and a junior engineer celebrates. The senior engineer checks recall on the fraud class: it is 0%. The model learned to output the mode.
The fix is to judge the model against that baseline with the right metrics — recall at a fixed false-alarm rate, or PR-AUC — and to use class weights or resampling during training.
For a bimodal case, think of delivery times from a city app. The histogram has two peaks, around 25 minutes and around 55 minutes. On inspection, the second peak is orders from outer suburbs. Two populations were mixed, and a single model feature "average delivery time" hid that. Adding a "zone" feature fixed it.
Follow-up questions to expect
- "Can the mode be used for numeric data?" — Yes for discrete counts (most common number of items in a cart), but for continuous values you bin them or use a density estimate and take the peak.
- "What does a bimodal distribution suggest?" — Usually two groups mixed together, such as weekday and weekend users or two device types. Split or add a feature that separates them.
- "Hard voting or soft voting?" — Soft voting averages probabilities and usually does better when the models are reasonably calibrated; hard voting takes the mode of predicted labels.