Course Content
Machine Learning Foundations
14 sections · 70 lessons
Can the same dataset be used for both regression and classification?
What you need to know
The target defines the task
The same rows of features can support many targets:
| Dataset | Regression target | Classification target |
|---|---|---|
| Flats in Bengaluru | Sale price in lakh | Budget, mid or premium |
| Loan applications | Expected loss in rupees | Default: yes or no |
| Telecom customers | Months until they leave | Will leave in 90 days: yes or no |
| Delivery orders | Minutes to deliver | Late (over 40 minutes): yes or no |
One dataset, two models
1import numpy as np2import pandas as pd3from sklearn.ensemble import RandomForestRegressor, RandomForestClassifier45rng = np.random.default_rng(42)6n = 10007df = pd.DataFrame({"area": rng.uniform(500, 2500, n), "dist_km": rng.uniform(2, 30, n)})8df["price"] = 20 + 0.07 * df.area - 1.2 * df.dist_km + rng.normal(0, 10, n) # lakh910# Same rows, two different targets11df["band"] = pd.cut(df.price, bins=[0, 60, 120, np.inf], labels=["budget", "mid", "premium"])12X = df[["area", "dist_km"]]1314reg = RandomForestRegressor(random_state=0).fit(X, df.price) # target: a number15clf = RandomForestClassifier(random_state=0).fit(X, df.band) # target: a category1617flat = pd.DataFrame({"area": [1150], "dist_km": [12]})18print("regression :", reg.predict(flat).round(1), "lakh")19print("classification:", clf.predict(flat), clf.predict_proba(flat).round(2))20print("class order :", clf.classes_)regression : [82.6] lakhclassification: ['mid'] [[0.03 0.97 0. ]]class order : ['budget' 'mid' 'premium']Same flat, same features. The regressor says 82.6 lakh. The classifier says "mid", with 97% confidence. pd.cut did the conversion from number to band.
What you lose when you bucket
The "mid" band runs from 60 to 120 lakh. A flat at 61 lakh and one at 119 lakh get the same label, and a flat at 59 lakh looks as different from the 61-lakh flat as a flat at 20 lakh does. The model can no longer learn that 61 is close to 59. So bucketing usually loses accuracy.
It is still the right call when:
- The decision really is categorical: approve or reject, send to manual review or not.
- The exact number is noisy or not trusted, but the band is.
- The users of the prediction act on bands, like "show premium listings to this user".
A useful middle path: predict the number, then bucket the prediction. You keep the full information in training and can move the cut-offs later without retraining.
Going from classes to numbers
Coding classes as 0, 1, 2 and running regression only makes sense when they are ordered and evenly spaced, such as a 1-to-5 star rating. Coding "Billing = 0, Network = 1, Sales = 2" as numbers invents an order that does not exist.
A real-life example
A lender first built a "default: yes or no" classifier. The finance team then asked, "How much will we lose?" A borrower who defaults after paying 11 of 12 EMIs costs far less than one who never pays. The team kept the classifier for approval decisions and added a regression model for the amount lost given default, trained on the same applications. The combination let them price loans by risk.
Follow-up questions to expect
- "Is predicting a star rating regression or classification?" — Either. Ratings are ordered, so regression (or ordinal regression) uses that order; classification ignores it. Try both and compare on the metric that matters.
- "How would you pick the band cut-offs?" — From the business, such as the price ranges the sales team already uses, not from what makes the model score higher.
- "Can one model do both?" — Yes, a neural network can have two heads, one for each output, trained together.