Machine Learning Foundations

Course Content

Machine Learning Foundations

14 sections · 70 lessons

Can the same dataset be used for both regression and classification?


What you need to know

The target defines the task

The same rows of features can support many targets:

DatasetRegression targetClassification target
Flats in BengaluruSale price in lakhBudget, mid or premium
Loan applicationsExpected loss in rupeesDefault: yes or no
Telecom customersMonths until they leaveWill leave in 90 days: yes or no
Delivery ordersMinutes to deliverLate (over 40 minutes): yes or no

One dataset, two models

Python
import numpy as npimport pandas as pdfrom sklearn.ensemble import RandomForestRegressor, RandomForestClassifierrng = np.random.default_rng(42)n = 1000df = pd.DataFrame({"area": rng.uniform(500, 2500, n), "dist_km": rng.uniform(2, 30, n)})df["price"] = 20 + 0.07 * df.area - 1.2 * df.dist_km + rng.normal(0, 10, n)   # lakh# Same rows, two different targetsdf["band"] = pd.cut(df.price, bins=[0, 60, 120, np.inf], labels=["budget", "mid", "premium"])X = df[["area", "dist_km"]]reg = RandomForestRegressor(random_state=0).fit(X, df.price)   # target: a numberclf = RandomForestClassifier(random_state=0).fit(X, df.band)   # target: a categoryflat = pd.DataFrame({"area": [1150], "dist_km": [12]})print("regression    :", reg.predict(flat).round(1), "lakh")print("classification:", clf.predict(flat), clf.predict_proba(flat).round(2))print("class order   :", clf.classes_)
Text
regression    : [82.6] lakhclassification: ['mid'] [[0.03 0.97 0.  ]]class order   : ['budget' 'mid' 'premium']

Same flat, same features. The regressor says 82.6 lakh. The classifier says "mid", with 97% confidence. pd.cut did the conversion from number to band.

What you lose when you bucket

The "mid" band runs from 60 to 120 lakh. A flat at 61 lakh and one at 119 lakh get the same label, and a flat at 59 lakh looks as different from the 61-lakh flat as a flat at 20 lakh does. The model can no longer learn that 61 is close to 59. So bucketing usually loses accuracy.

It is still the right call when:

  • The decision really is categorical: approve or reject, send to manual review or not.
  • The exact number is noisy or not trusted, but the band is.
  • The users of the prediction act on bands, like "show premium listings to this user".

A useful middle path: predict the number, then bucket the prediction. You keep the full information in training and can move the cut-offs later without retraining.

Going from classes to numbers

Coding classes as 0, 1, 2 and running regression only makes sense when they are ordered and evenly spaced, such as a 1-to-5 star rating. Coding "Billing = 0, Network = 1, Sales = 2" as numbers invents an order that does not exist.

A real-life example

A lender first built a "default: yes or no" classifier. The finance team then asked, "How much will we lose?" A borrower who defaults after paying 11 of 12 EMIs costs far less than one who never pays. The team kept the classifier for approval decisions and added a regression model for the amount lost given default, trained on the same applications. The combination let them price loans by risk.

Follow-up questions to expect

  • "Is predicting a star rating regression or classification?" — Either. Ratings are ordered, so regression (or ordinal regression) uses that order; classification ignores it. Try both and compare on the metric that matters.
  • "How would you pick the band cut-offs?" — From the business, such as the price ranges the sales team already uses, not from what makes the model score higher.
  • "Can one model do both?" — Yes, a neural network can have two heads, one for each output, trained together.