Course Content
Introduction to AI
4 sections · 10 lessons
Key Components — Data, Algorithms, Evaluation
Every AI system rests on three things: the data it learns from, the algorithm that does the learning, and the evaluation that tells you whether it worked.
Newcomers spend nearly all their attention on the middle one. Practitioners will tell you the first and third matter more. This lesson explains why, and gives you enough grounding in each to reason about real systems rather than repeat slogans.
Data: the part that sets the ceiling
There is a claim worth stating plainly at the start:
Better data beats a better algorithm, nearly every time. Given a fixed budget, improving the data usually buys more than switching to a more sophisticated model. Practitioners with a decade of experience say this repeatedly, and beginners consistently do not believe it until they have been burned.
The reason is structural. An algorithm can only find patterns that exist in the data it is shown. If the signal is not there, no amount of modelling sophistication conjures it. If the data is misleading, sophistication just learns the misleading pattern more efficiently.
What "good data" means, concretely
Relevant. It contains information that actually relates to what you are predicting. Predicting equipment failure from purchase date alone will fail if failure depends on usage intensity you never recorded.
Representative. It resembles what the system will meet in the real world. This is the property most often violated and most often fatal. A voice model trained on studio recordings meets people in cars, kitchens and on the street.
Sufficient in variety. Enough range to cover the situations that will occur. Not just volume — a million photographs all taken at noon do not teach a model about dusk.
Accurately labelled. The answers attached to your examples are correct. Sloppy labels teach confident mistakes.
Recent enough. The relationship you are learning still holds. Consumer behaviour data from before a major disruption may describe a world that no longer exists.
Features: what the model actually sees
A model does not see a customer. It sees numbers. Deciding which numbers is called feature engineering, and on tabular data it is frequently where the real gains are.
Suppose you are predicting whether a customer will cancel a subscription. The raw data has a signup date and a list of logins. Neither is directly useful. What the model can learn from:
- Days since signup
- Logins in the last 30 days
- Change in login frequency versus the previous 30 days
- Days since last login
- Number of support tickets raised
That third one — the change in behaviour — is often far more predictive than any absolute number. Someone logging in twice a week is fine; someone who dropped from daily to twice a week is leaving. The raw data contained this, but only as an implication. Someone had to make it explicit.
This is the essence of feature engineering: taking what you know about the problem and making it visible to a system that has no world knowledge of its own. Deep learning reduces the need for this on images, audio and text — but on business data in tables, it remains where much of the value is.
The ways data goes wrong
Sampling bias. Your data over-represents some groups. A model trained on app users learns about people comfortable with apps, then gets deployed to everyone.
Historical bias. The data faithfully records past human decisions, including unfair ones. The model reproduces them, at scale, wearing a mathematical costume.
Survivorship bias. You only have data on things that survived to be recorded. Analysing successful companies to find what makes companies succeed misses every company that did the same things and failed.
Label noise. Answers are wrong or inconsistent. Two annotators applying different interpretations of the same guideline produce a target the model cannot hit.
Algorithms: choosing how to learn
You do not need to implement these. You need to know roughly what each does, what it costs, and when it is the wrong choice.
The main families
| Family | How it works | Good at | Weak at |
|---|---|---|---|
| Linear models | Fits a weighted sum of inputs | Fast, fully explainable, works on small data | Cannot capture complex interactions |
| Decision trees | A series of yes/no splits | Readable, handles mixed data types | A single tree overfits badly |
| Tree ensembles | Many trees combined | Usually the best choice for tabular data | Harder to explain than one tree |
| Neural networks | Layers of learned transformations | Images, audio, text, very large data | Data-hungry, expensive, opaque |
| Clustering | Groups similar items | Finding structure with no labels | You must interpret the groups yourself |
The row worth remembering: for data in a table, tree ensembles usually beat neural networks. This surprises people who assume deep learning is uniformly superior. On spreadsheet-shaped problems, ensembles typically win on accuracy, train in seconds rather than hours, and need far less data.
How to choose
A sequence that will serve you well:
- Start with the simplest thing that could work. A linear model or a single tree. It trains in seconds and gives you a baseline.
- Check whether that is already good enough. Sometimes it is, and you have saved weeks.
- Move up only if the baseline falls short, and measure whether the gain justifies the cost.
- Let the data type decide the ceiling. Tables → ensembles. Images, audio, text → neural networks.
Two constraints often matter more than accuracy, and are frequently forgotten until late:
Explainability. If you must legally tell someone why they were refused, an opaque model is unusable no matter how accurate.
Cost and latency. A model that takes two seconds to respond is unusable in a checkout flow, whatever its accuracy.
Parameters the model learns, and settings you choose
A distinction that confuses people and is genuinely simple.
Parameters are learned during training — the weights the model adjusts to fit the data. You never set these by hand.
Hyperparameters are settings you choose before training: how deep a tree may grow, how many layers a network has, how large a step training takes. The model cannot learn these; you search for good values by trying combinations and comparing on the validation set.
This is exactly why the validation set exists as something separate from the test set. Choosing hyperparameters by looking at results is a form of fitting — so you need a third, untouched set to get an honest final number.
Evaluation: knowing whether it actually worked
The third pillar, and the one where self-deception is easiest.
Reading a confusion matrix
Every prediction on a yes/no problem falls into one of four boxes. Learn these and most evaluation becomes obvious.
| Model says yes | Model says no | |
|---|---|---|
| Actually yes | True positive — correctly caught | False negative — missed it |
| Actually no | False positive — false alarm | True negative — correctly ignored |
Everything else is arithmetic on these four numbers.
Take a fraud detector run over 10,000 transactions containing 100 real frauds. It flags 150 transactions; 80 are genuine fraud.
- Precision = 80 / 150 = 53%. Of what it flagged, about half were real. Nearly half of flagged customers were inconvenienced for nothing.
- Recall = 80 / 100 = 80%. Of real frauds, it caught four in five. Twenty got through.
- Accuracy = 9,860 / 10,000 = 98.6% — a number that tells you almost nothing useful here.
Look at how differently those three describe the same system. This is why a single headline number should always make you ask which one.
The threshold is a business decision
Most models output a probability, and someone chooses the cut-off at which action is taken. Move it and you slide along the trade:
- Lower threshold → flag more → recall rises, precision falls → catch more fraud, annoy more customers.
- Higher threshold → flag less → precision rises, recall falls → fewer false alarms, more fraud gets through.
The model does not decide this. A human weighs the cost of each error type and picks. Understanding that this dial exists — and that it is not a technical setting — is one of the more practically useful things in this lesson.
When the answer is a number
For predicting quantities rather than categories, two measures dominate:
Mean absolute error — the average size of the mistake, in the original units. If you are predicting house prices and this is £15,000, you are typically off by £15,000. Easy to explain to anyone.
Root mean squared error — similar, but squaring the errors first means large mistakes count disproportionately. Use it when occasional big misses are much worse than consistent small ones.
Choose based on which kind of error actually hurts. If being wrong by £100,000 once is far worse than being wrong by £10,000 ten times, the squared measure reflects your situation better.
Ways evaluation lies to you
Testing on training data. Guarantees excellent, meaningless numbers.
Reusing the test set repeatedly. Each time you adjust based on test results, your estimate becomes more optimistic. In effect you are slowly fitting to the test set.
Ignoring subgroups. A model that is 90% accurate overall might be 95% for one group and 60% for another. The average conceals the failure. Always evaluate on the groups your system will actually affect.
Evaluating on stale data. If the world moved since your test set was collected, your numbers describe a world that no longer exists.
How the three fit together
The pillars are not independent, and the dependencies run in one direction.
- Data sets the ceiling. No algorithm extracts a pattern that is not present.
- The algorithm determines how close to that ceiling you get. A better method helps — but only up to what the data supports.
- Evaluation tells you where you actually are. Get it wrong and you cannot tell success from failure, which makes the other two unimprovable.
Which leads to the practical priority most teams get backwards. Faced with a model that is not good enough, the instinct is to try a more sophisticated algorithm. The higher-yield moves, in order, are usually: check your evaluation is honest, then improve the data, then engineer better features, and only then reach for a better algorithm.
Check your understanding
0 of 3 answered
1.A fraud model flags 200 transactions. 50 of them are real fraud, and there were 100 real frauds in total. What are its precision and recall?
2.The bank complains that too many genuine customers are having payments declined. Which change to the threshold helps, and what does it cost?
3.A team's model on tabular customer data is not good enough. According to this lesson, what should they check first?