Course Content
Machine Learning Foundations
14 sections · 70 lessons
What is regression in Machine Learning?
What you need to know
What "continuous" means
A regression target is a number on a scale where the distance between values means something. A flat priced at 90 lakh is 10 lakh more than one at 80 lakh, and being off by 2 lakh is better than being off by 20. That is the test: if "how far off" makes sense, it is regression.
Predicting Bengaluru flat prices
1import numpy as np2from sklearn.linear_model import LinearRegression3from sklearn.model_selection import train_test_split4from sklearn.metrics import mean_absolute_error, root_mean_squared_error, r2_score56rng = np.random.default_rng(42)7n = 10008area = rng.uniform(500, 2500, n) # sq ft9dist = rng.uniform(2, 30, n) # km from MG Road10age = rng.uniform(0, 25, n) # building age in years11price = 20 + 0.07 * area - 1.2 * dist - 0.5 * age + rng.normal(0, 10, n) # lakh1213X = np.column_stack([area, dist, age])14X_tr, X_te, y_tr, y_te = train_test_split(X, price, test_size=0.2, random_state=0)15pred = LinearRegression().fit(X_tr, y_tr).predict(X_te)1617print("MAE :", round(mean_absolute_error(y_te, pred), 2), "lakh")18print("RMSE:", round(root_mean_squared_error(y_te, pred), 2), "lakh")19print("R2 :", round(r2_score(y_te, pred), 3))MAE : 7.41 lakhRMSE: 8.94 lakhR2 : 0.955The data is simulated, with random noise of about 10 lakh built in, so no model can do much better. Read the output like this: on flats it never saw, the model is off by 7.41 lakh on average; RMSE is a little higher because it weighs the bigger misses more; and it explains 95.5% of the variation in price.
The three metrics in plain words
| Metric | What it means | Use it when |
|---|---|---|
| MAE (mean absolute error) | Average size of the error, in rupees, minutes, units | You want a number a manager understands; all errors cost the same per unit |
| RMSE (root mean squared error) | Squares errors before averaging, so big misses count much more | One huge error is much worse than several small ones |
| R² | Share of the target's variation the model explains; 1.0 is perfect, 0 is no better than predicting the mean | Comparing models on the same data |
Algorithms
Linear regression is the baseline: fast, explainable, a good first model. Tree-based models (random forest, gradient boosting such as XGBoost or LightGBM) usually win on tabular data because they capture non-linear effects, such as price per square foot jumping near a metro station.
A real-life example
A food-delivery app predicts delivery time in minutes. MAE is 4.2 minutes and looks fine. But RMSE is 9.8 minutes, far higher, which means some orders are off by 20 to 30 minutes. The team finds these are all rainy evenings in two zones. Customers forgive a 4-minute miss but cancel after a 25-minute one, so the team optimises RMSE and adds a rain feature. MAE barely moves; complaints drop.
Follow-up questions to expect
- "What is the difference between MAE and RMSE?" — MAE averages absolute errors, so every rupee of error counts equally. RMSE squares errors first, so a few big misses raise it sharply. RMSE is always greater than or equal to MAE.
- "Can R² be negative?" — Yes, on test data. It means the model is worse than simply predicting the average of the target.
- "Is logistic regression a regression model?" — Despite the name, it is used for classification. It predicts a probability, which is then turned into a class.