Machine Learning Foundations

Course Content

Machine Learning Foundations

14 sections · 70 lessons

What is Machine Learning, and how is it different from traditional programming?


Who writes the spam ruleTraditional programming• A developer listswords like lottery, free• Data plus rules go in, answers come out• Misses the KYC scam: no keyword matched• Exact, auditable, cheap to runMachine learning• Eight labelled messages go in• Data plus answers go in, rules come out• Catches the KYC scam from similar words• Needs data, and is sometimes wrong
Nobody wrote the KYC rule — the labelled examples did, which is the whole difference in one line.

What you need to know

Two ways to get a program to make a decision

Suppose a bank wants software that says "approve" or "reject" for a loan application.

  • Traditional programming. An analyst writes the rule: if income > 50000 and existing_loans < 2: approve. The computer follows it exactly. If the rule is wrong, a human must notice and change it.
  • Machine learning. You collect 50,000 past loans with the outcome you care about (did the borrower repay or default?). A learning algorithm studies them and finds the pattern that links the inputs (income, age, existing loans, repayment history) to the outcome. The pattern it finds is the model.
Text
Traditional programming:  data + rules   -> answersMachine learning:         data + answers -> rules (the model)

The second line is the whole idea. After training, you use the model like any function: new application in, prediction out.

A small example you can run

Here is the same problem, spam detection, solved both ways with scikit-learn.

Python
from sklearn.feature_extraction.text import CountVectorizerfrom sklearn.naive_bayes import MultinomialNB# Traditional programming: a human writes the ruledef is_spam_rule(text):    return any(w in text.lower() for w in ["lottery", "winner", "free"])# Machine learning: we give examples with answers, the algorithm finds the rulemessages = [    "You are a lottery winner, claim now", "Free recharge offer, click link",    "Get a free iPhone today", "Urgent: your KYC expires, update now",    "Lunch at 1?", "Meeting moved to 4 pm", "Can you share the report",    "Your OTP for login is 4821, do not share",]labels = [1, 1, 1, 1, 0, 0, 0, 0]          # 1 = spam, 0 = not spamvec = CountVectorizer()model = MultinomialNB().fit(vec.fit_transform(messages), labels)new = ["Update KYC now or account blocked", "Is the free parking open?"]print("rule :", [is_spam_rule(m) for m in new])print("model:", model.predict(vec.transform(new)).tolist())
Text
rule : [False, True]model: [1, 0]

The hand-written rule misses the KYC scam (no keyword matched) and wrongly flags an innocent message that contains "free". The model learned from the examples that "KYC", "update" and "now" appear in spam, so it gets both right. With eight messages this is a toy, but the point holds at scale: nobody wrote the KYC rule, the data did.

When ML is the right tool, and when it is not

Use machine learning when...Use plain code when...
The rules are too many to write (spam, fraud, image recognition)A few clear rules cover it (GST calculation, age check)
The pattern changes over time (new scam styles)The logic is fixed by law or policy
You have many past examples with known outcomesYou have no data, or very little
Small errors are acceptableEvery decision must be exactly right and explainable

ML has costs that plain code does not: you need data, the model is sometimes wrong, and it can quietly get worse when the world changes. So it is a trade, not an upgrade.

A real-life example

Production scenario. A UPI payments app first blocked fraud with 40 hand-written rules such as "block if more than 10 transfers in 5 minutes". Fraudsters learned the rules within weeks and stayed just under each limit. The team trained a model on two years of transactions labelled fraud or genuine, using hundreds of signals together. It caught patterns no single rule described. They kept a few hard rules on top, like "never allow a transfer to a blocked account", because some decisions must be certain.

Follow-up questions to expect

  • "Can you give an example where you would not use ML?" — Calculating income tax. The rules are exact, written by law, and must be 100% correct, so code is cheaper, faster and auditable.
  • "What does the model actually learn?" — Numbers called parameters, such as the weights in a linear model or the split points in a decision tree, chosen so predictions match the training answers as closely as possible.
  • "Is ML always better than rules?" — No. Many production systems combine both: a model scores the risk, and rules handle hard limits and legal requirements.