Machine Learning Foundations

Course Content

Machine Learning Foundations

14 sections · 70 lessons

What does a “model” mean in Machine Learning?


What you need to know

Algorithm, model, prediction

These three words are often mixed up. Keep them apart:

  • Algorithm — the recipe for learning, such as "linear regression" or "random forest". It is code that exists before you have any data.
  • Model — what the algorithm produces after it sees your data. It holds the learned numbers, called parameters.
  • Prediction — the model's output for one new input.

A model is just learned numbers

Here is a model that prices flats in Bengaluru from area and number of bedrooms.

Python
import joblibfrom sklearn.linear_model import LinearRegression# Bengaluru flats: [area in sq ft, bedrooms] -> price in lakh rupeesX = [[650, 1], [900, 2], [1100, 2], [1250, 2], [1400, 3], [1650, 3], [1900, 3], [2300, 4]]y = [48, 66, 80, 88, 101, 116, 133, 160]model = LinearRegression().fit(X, y)          # the algorithm runs hereprint("coef_     :", model.coef_.round(3))    # the learned parametersprint("intercept_:", round(model.intercept_, 2))print("prediction for 1200 sq ft, 2 BHK:", model.predict([[1200, 2]]).round(1))joblib.dump(model, "house_model.joblib")     # the artefact you deploy
Text
coef_     : [0.064 2.43 ]intercept_: 4.3prediction for 1200 sq ft, 2 BHK: [85.5]

The whole "intelligence" of this model is three numbers. It says: price in lakhs = 4.3 + 0.0636 × area + 2.43 × bedrooms (the first coefficient prints as 0.064 after rounding). For 1,200 sq ft and 2 bedrooms that is 4.3 + 76.3 + 4.9 = 85.5 lakh. fit is where learning happens; predict just does the arithmetic.

What the model looks like for other algorithms

AlgorithmWhat the trained model stores
Linear or logistic regressionOne weight per feature, plus an intercept
Decision treeA tree of questions like "area greater than 1,300?" and a value at each leaf
Random forestHundreds of such trees; the prediction is their average or vote
Neural networkWeight matrices, from thousands to billions of numbers

The model in production

In a real system the model is a file, such as a .joblib, ONNX or .safetensors file. It gets a version number, is loaded by a prediction service, is monitored, and can be rolled back like any release. When someone says "we deployed model v7", they mean this file.

A real-life example

Production scenario. A property portal retrains its price model every month. In March, v12 suddenly predicted flats in Whitefield 20% too high. Because each model file was versioned and stored with the data it was trained on, the team rolled back to v11 in minutes, then found that a scraping bug had added duplicate luxury listings to the March training data.

Follow-up questions to expect

  • "What is the difference between a parameter and a hyperparameter?" — Parameters are learned from data, such as regression weights. Hyperparameters are chosen by you before training, such as tree depth or learning rate.
  • "Is a pipeline part of the model?" — In practice, yes. The preprocessing steps (scaling, encoding) must be saved with the model, or predictions will be wrong. scikit-learn's Pipeline bundles them into one object.
  • "Is a bigger model always better?" — No. More parameters can memorise noise (overfitting) and cost more to serve.