Course Content
Introduction to AI
4 sections · 10 lessons
AI vs ML vs Deep Learning vs Generative AI
Four terms get thrown around as if they were interchangeable: artificial intelligence, machine learning, deep learning, generative AI. Marketing departments use whichever sounds most impressive. News articles switch between them mid-paragraph.
They are not synonyms. They describe four different things at four different levels of specificity, and mixing them up will cost you — in interviews, in technical conversations, and most expensively when choosing a tool for a real project.
This lesson fixes that permanently. By the end you will be able to hear any AI claim and place it correctly, and you will know which of the four you actually need for a given problem.
The one picture that makes it click
These four terms are not four categories sitting side by side. They are nested inside one another, like a set of bowls stacked in a cupboard.
┌───────────────────────────────────────────────────────┐│ ARTIFICIAL INTELLIGENCE ││ Any software doing tasks that need judgement ││ ││ ┌───────────────────────────────────────────────┐ ││ │ MACHINE LEARNING │ ││ │ Learns from examples instead of hand rules │ ││ │ │ ││ │ ┌───────────────────────────────────────┐ │ ││ │ │ DEEP LEARNING │ │ ││ │ │ Many-layered neural networks │ │ ││ │ │ │ │ ││ │ │ ┌───────────────────────────────┐ │ │ ││ │ │ │ GENERATIVE AI │ │ │ ││ │ │ │ Creates new content │ │ │ ││ │ │ └───────────────────────────────┘ │ │ ││ │ └───────────────────────────────────────┘ │ ││ └───────────────────────────────────────────────┘ │└───────────────────────────────────────────────────────┘Read it inward and each statement is true:
- Every generative AI system is deep learning.
- Every deep learning system is machine learning.
- Every machine learning system is AI.
Read it outward and each statement is false:
- Not every AI system is machine learning.
- Not every machine learning system is deep learning.
- Not every deep learning system is generative.
The nesting runs one way only. Almost every mistake people make with these words comes from reading the arrow backwards.
Artificial intelligence — the outermost bowl
AI is any software that performs tasks normally requiring human judgement. That is the whole definition. It says nothing about how.
This is why AI is a genuinely old field. A chess program written in 1985, using move-scoring rules a grandmaster dictated, is AI. It plays a game that needs judgement. It contains no learning whatsoever — it knows exactly what it was told and will know exactly that forever.
Things that are AI but not machine learning:
- A thermostat schedule built from hand-written
ifconditions. - A route planner using a classical shortest-path algorithm.
- A medical expert system built from interviewing doctors.
- The enemy behaviour in most video games — usually decision trees a designer wrote.
These are sometimes called symbolic AI or good old-fashioned AI. They are not obsolete. Rule-based systems are still the right answer when rules exist, because they are cheaper, faster, fully explainable, and cannot surprise you.
Machine learning — learning the rules from examples
Machine learning is the subset of AI where the system derives its own rules from data.
You supply examples and correct answers. The system finds the pattern. Nobody writes the logic.
The three families of machine learning
Almost every ML system fits one of three shapes, distinguished by what kind of feedback the system gets.
Supervised learning — learning from labelled answers
You give the system inputs and the correct output for each. It learns the mapping.
1# Each house has features, and a price we already know.2features = [[1200, 3, 1995], [1800, 4, 2010], [950, 2, 1988]]3prices = [250000, 410000, 180000]45model.fit(features, prices)6model.predict([[1500, 3, 2005]]) # -> 320000This is by far the most common family in commercial use. Spam detection, price prediction, medical image screening, credit scoring, demand forecasting — all supervised.
Its cost is obvious once you see it: somebody had to label every example. Ten thousand labelled X-rays means a radiologist spent months. Labelling is often the single largest expense in a machine learning project, and the quality of those labels caps everything.
Unsupervised learning — finding structure with no answers given
You give the system data with no labels and ask it to find structure.
# No correct answers supplied. Just: find natural groups.model.fit(customer_data)model.predict(new_customer) # -> group 3The system might discover that your customers fall into five natural groups. It cannot tell you those groups are "bargain hunters" and "weekend browsers" — a human has to look at each group and interpret it. The system finds structure; meaning is your job.
Used for customer segmentation, anomaly detection, and compressing complicated data into something manageable.
Reinforcement learning — learning from consequences
No labelled answers. The system tries things, receives reward or penalty, and learns a strategy that earns more reward.
This is how a system learns to play a game with no one showing it good moves — it plays enormous numbers of games and notices which choices tend to precede winning. Also used in robotics and in tuning conversational models to be more helpful.
It is powerful and it is difficult, because the reward has to be defined by a human, and systems reliably find loopholes. Reward a cleaning robot for collected dust and it may learn to tip the bin out and collect it again. It did exactly what you asked. You asked for the wrong thing.
Choosing between them
| Supervised | Unsupervised | Reinforcement | |
|---|---|---|---|
| You provide | Inputs and correct answers | Inputs only | An environment and a reward |
| It learns | Input → output mapping | Hidden structure | A strategy |
| Main cost | Labelling | Interpreting results | Defining reward correctly |
| Typical use | Prediction, classification | Segmentation, anomalies | Games, robotics, control |
Deep learning — machine learning with many layers
Deep learning is the subset of machine learning that uses neural networks with many layers. "Deep" refers to the number of layers, nothing more.
What a layer actually does
Ignore the biological metaphors; they cause more confusion than they resolve. Practically, each layer takes numbers in, multiplies them by weights it has learned, and passes numbers out. Stack many layers and each one builds on what the previous one found.
For an image recognition network, this progression is real and has been visualised:
- Early layers detect edges and simple colour boundaries.
- Middle layers combine edges into shapes — corners, curves, textures.
- Later layers combine shapes into parts — an eye, a wheel, a leaf.
- Final layers combine parts into whole objects — a face, a car, a tree.
Nobody programmed that hierarchy. Nobody wrote "look for edges first". It emerged because that happens to be an efficient way to organise the problem, and the training process found it.
The one thing deep learning changed
Before deep learning, a large part of a machine learning practitioner's job was feature engineering: deciding by hand which measurable properties of the input the model should look at.
Building a face detector meant a human deciding to measure eye spacing, nose width, jaw curvature — and if those were the wrong properties, no algorithm could save you.
Deep learning largely removed that step. Give a network raw pixels and enough examples, and it works out which features matter. That is the reason for the leap in image and speech performance after 2012.
What it costs you
| Classical ML | Deep learning | |
|---|---|---|
| Data needed | Hundreds to thousands | Tens of thousands upward |
| Feature engineering | Substantial human work | Mostly automatic |
| Training cost | Minutes on a laptop | Hours to weeks on specialised hardware |
| Explainability | Often readable | Very difficult |
| Best for | Tables, modest data | Images, audio, text, large data |
This table contains a practical warning. Deep learning is not automatically better. On a spreadsheet of ten thousand rows, a classical method will often beat a neural network while training in seconds and remaining explainable. Reaching for deep learning by default is a common and costly instinct.
Generative AI — producing rather than labelling
Generative AI is deep learning that creates new content instead of selecting an answer.
The distinction is in the shape of the output.
| Discriminative (traditional) | Generative | |
|---|---|---|
| Question it answers | Which category is this? | What would plausibly come next? |
| Output | One of a fixed set | New content, unlimited possibilities |
| Example | "This photo is a cat" (0.94) | Produces a picture of a cat |
| Same input twice | Same answer | Often a different answer |
That last row surprises people and is worth understanding. Generative models sample from a probability distribution rather than picking a maximum. Ask the same question twice and you may get two different valid answers — by design, not by fault.
How a language model actually works
The mechanism is simpler than most people expect, and knowing it explains most of the behaviour you will encounter.
A language model does exactly one thing: given some text, predict what comes next. That is it. Everything else is a consequence.
Give it "The capital of France is" and it produces a probability for every possible next word — perhaps 0.91 for "Paris", 0.02 for "a", 0.01 for "located". It picks one, adds it to the text, and repeats.
Now hold two facts together:
- The model is only ever predicting a plausible next word.
- It has been trained on a very large fraction of everything humans have written down.
To predict text well across all of that, it has to internalise a great deal — grammar, facts, argument structure, code syntax, conversational conventions. Capability came out as a side effect of doing one narrow task extremely well.
The chat assistants you use get further training on top of this: on examples of helpful answers, on people's ratings of which reply was better, and, for reasoning models, on rewards for reaching correct solutions. That shapes how they respond. Underneath, each word is still produced by the same next-word prediction.
Why they invent things
This directly explains the failure mode people find most baffling. When a model states a confident falsehood — a fake citation, an invented statistic — it is not malfunctioning. It is doing precisely what it was built to do.
The model is optimising for plausible, not for true. A fabricated citation is plausible text: it has the right shape, the right kind of author name, a believable journal. Nothing in the mechanism checks whether the thing exists.
This is why "hallucination" is misleading as a term. It suggests a glitch. It is better understood as the system's normal behaviour showing through in a case where plausibility and truth come apart. Which is exactly why you verify anything factual a model tells you.
Connecting a model to web search or to your own documents, which you will meet in the next section, reduces invention because the answer can come from a real source. It does not end it: the model can still misread or misquote what it was given.
Choosing the right one
Work down this list and stop at the first match.
- Can you write the rules? → Use ordinary software. Do not use AI. It is cheaper, faster, and fully predictable.
- Do you have labelled examples and want a prediction or a category? → Supervised machine learning. Start with a classical method.
- Is your data images, audio, or free text, and do you have a lot of it? → Deep learning.
- Do you need new content produced — text, images, code? → Generative AI.
- Do you have data but no labels and want to discover structure? → Unsupervised learning.
Two mistakes are especially common and especially expensive:
Using generative AI for classification. Asking a large language model "is this review positive or negative?" works, but it is slower, more expensive, and less consistent than a small purpose-trained classifier — sometimes by a factor of a hundred in cost.
Using deep learning on small tabular data. With a few thousand spreadsheet rows, classical methods usually win on accuracy and train in seconds and let you explain the result.
Reading claims critically
"AI-powered" on a product page means almost nothing. It could be a large language model, or it could be a single if statement.
Useful questions to ask instead:
- Which of the four is it — rules, classical ML, deep learning, or generative?
- What was it trained on, and does that data resemble my situation?
- Does it output a category or new content?
- What is its accuracy, and what happens when it is wrong?
A vendor who cannot answer these is either not technical or not being straight with you. Either way you have learned something useful.
Check your understanding
0 of 3 answered
1.A system is given a year of customer purchase records with no labels and asked to find natural groups of customers. Which kind of machine learning is this?
2.You have 5,000 rows of customer data in a spreadsheet and want to predict who will cancel next month. What should you try first?
3.A language model gives a confident reference to a research paper that does not exist. What is the best explanation?