Course Content
Introduction to AI
4 sections · 10 lessons
AI Pipeline — Data to Deployment
Ask someone what building an AI system involves and they will describe training a model. Training is one step, and it is rarely the hard one.
A working AI system is a pipeline — a sequence of stages that carries raw information from the world through to a decision that affects someone, and then back around again. Most projects that fail do not fail at the modelling stage. They fail at the boring parts on either side of it.
This lesson walks the whole pipeline, stage by stage, with the mistakes that kill projects at each one.
1. Define → What decision are we changing? 2. Collect → Get the data 3. Prepare → Clean it, label it, split it 4. Train → Fit a model 5. Evaluate → Is it good enough, and by what measure? 6. Deploy → Make it available where it is needed 7. Monitor → Watch it degrade └────────────────── back to 2Note the arrow at the bottom. This is a loop, not a line. A model deployed and forgotten is a model quietly getting worse.
Stage 1 — Define the problem
The cheapest stage to do well and the most expensive to get wrong, because every later stage inherits its mistakes.
Three questions have to be answered before anyone writes code.
What decision changes as a result of this? Not "we want to use AI for customer service" — that is a wish. "We want to route incoming tickets to the right team automatically, because manual routing takes four hours and delays resolution." Now there is a decision, a current cost, and something to measure.
What does the model output, exactly? A category from a fixed list? A number? A ranked order? A piece of text? This determines everything downstream — what data you need, what kind of model, how you evaluate it. Vagueness here compounds.
What is good enough? Decide the threshold before you build, not after. This forces the honest conversation: what happens when it is wrong, who catches that, and what does a mistake cost? A team that has not answered this will accept whatever accuracy they happen to achieve and call it a success.
The most valuable question at this stage is the one people skip: could a simple rule do this? If a handful of conditions gets you 80% of the value, build that first. It ships in a week, it is fully explainable, and it gives you a baseline the model must beat to justify its cost.
Stage 2 — Collect the data
Data comes from somewhere, and where it comes from determines what the model can ever learn.
Typical sources: your own systems (transaction records, logs, support tickets), sensors, public datasets, purchased data, or data you commission from scratch.
The mistake that ruins projects here
The training data does not resemble the situation where the system will run. This is the single most common cause of a model that tests brilliantly and fails in production.
Examples of the shape it takes:
- Training on photographs taken in good light, deploying to a warehouse lit by fluorescent tubes.
- Training on data from your largest market, deploying globally.
- Training on last year's customers, deploying after your product and pricing changed.
- Training on data from one factory's sensors, deploying to a factory with different equipment.
Ask early and bluntly: in what specific ways will the live data differ from what I am training on? Write the answer down. It will predict most of your future problems.
How much do you need?
There is no universal number, but there are useful anchors: hundreds to low thousands of examples for a simple classical model on tabular data; tens of thousands upward for deep learning on images or text; and if you are adapting an existing large model rather than training from scratch, sometimes only hundreds.
Variety usually matters more than raw volume. Ten thousand near-identical examples teach less than one thousand genuinely varied ones.
Stage 3 — Prepare the data
This routinely consumes the majority of a project's time. Nobody enjoys it and everybody underestimates it.
Cleaning
Real data is broken in predictable ways: missing values, duplicates, inconsistent formats, obvious errors, and outliers that may be mistakes or may be exactly what you are trying to detect.
Each requires a decision, and each decision has consequences. Consider missing values. You can drop those rows — but if data is missing for a reason (customers who declined to state income might differ systematically), you have just introduced bias. You can fill them with an average — but now you have invented data that will look confidently normal to the model. There is no universally right answer, only a choice you should make deliberately and document.
Labelling
For supervised learning, someone must attach the correct answer to each example. This is expensive and its quality caps everything.
Two problems arise reliably. Labellers disagree with each other — measure this deliberately, because if two experts agree only 70% of the time, no model will exceed that ceiling and you have learned something important about the task. And labels drift as guidelines are interpreted differently over months, leaving the model learning from a moving target.
Splitting — and the mistake that invalidates everything
You divide your data into three parts:
| Set | Share | Purpose |
|---|---|---|
| Training | ~70% | The model learns from this |
| Validation | ~15% | Tune settings and compare approaches |
| Test | ~15% | Final honest estimate — touched once |
The rule that makes this meaningful: the test set must be untouched until the very end. Every time you look at test results and adjust something, you leak information about it into your choices, and your estimate becomes optimistic.
Now the error that silently ruins results — data leakage. It happens when information from the test set influences training, and it produces excellent numbers that mean nothing.
Two common forms:
- Splitting time-series randomly. If you are predicting the future, a random split lets the model train on Thursday and be tested on Wednesday. It has effectively seen the answer. Always split by time for time-dependent problems.
- Computing statistics before splitting. Scaling your features using the average of the whole dataset leaks information about the test set into training. Compute on training data only, then apply those same values to validation and test.
If a model's results look surprisingly good, suspect leakage before celebrating. Experienced practitioners treat an unexpectedly high score as a bug report, not a success.
Stage 4 — Train the model
Ironically the most automated stage. Conceptually the loop is simple:
for each round: predictions = model(training_data) error = compare(predictions, correct_answers) adjust_model_to_reduce(error)Repeat until the error stops improving.
The one concept you must understand: overfitting
This is the central failure mode of machine learning, and the intuition matters more than the maths.
A model can reduce its training error by learning genuine patterns — or by memorising the training examples, including their noise and accidents. The second gets a perfect training score and fails completely on anything new.
The analogy that captures it: a student who memorises past exam papers word for word scores perfectly on those papers and cannot answer a new question. They learned the answers, not the subject.
How you detect it:
| Training error | Validation error | Diagnosis |
|---|---|---|
| Low | Low | Working well |
| Low | High | Overfitting — memorising |
| High | High | Underfitting — too simple, or bad features |
That table is the reason the validation set exists. Training error alone tells you nothing about whether the model has learned anything transferable.
The remedies: more and more varied data; a simpler model; techniques that penalise complexity; and stopping training at the point where validation error starts rising even though training error keeps falling.
Stage 5 — Evaluate
The stage where honest projects diverge from ones that are fooling themselves.
Why accuracy is usually the wrong measure
Take a disease affecting 1 in 1,000 people. A model that answers "healthy" for everyone is 99.9% accurate and detects nothing.
Accuracy collapses whenever one outcome is much rarer than the other — which describes most problems worth solving. You need to separate the two kinds of error:
| Measure | The question it answers | Matters most when |
|---|---|---|
| Precision | Of the cases we flagged, how many were real? | False alarms are costly |
| Recall | Of the real cases, how many did we catch? | Missing one is costly |
These trade against each other, always. Flag more aggressively and you catch more real cases while raising false alarms. The right balance is a business decision, not a technical one.
Cancer screening leans toward recall — a false alarm means an unnecessary follow-up test, while a miss can be fatal. A spam filter leans toward precision — some spam getting through is annoying, but deleting a real email is unacceptable.
Always compare against a baseline
"Our model is 87% accurate" is meaningless alone. Compared to what?
- Always predicting the most common answer
- A simple rule anyone could write in an afternoon
- Whatever the business currently does
If a three-line rule gets 85%, your 87% model may not justify its cost, complexity, and ongoing maintenance. This comparison is skipped remarkably often, and it is where a lot of unnecessary machine learning comes from.
Stage 6 — Deploy
Moving from a model that works on your machine to one that serves real users. The choice is usually between two shapes:
Batch — run periodically over many records and store the results. Simple, cheap, easy to debug. Suits anything that does not need to be instant, like nightly risk scoring.
Real time — serve behind an API and respond in milliseconds. Necessary for fraud checks at the till, or anything a user waits on. More complex and more expensive.
Practical concerns that surprise people: response time budgets are tight and models are slow; cost scales with traffic in ways that do not show up in testing; you need model versioning so you can roll back; and you should release gradually — a small share of traffic first — because the first real contact with production data is where surprises live.
Stage 7 — Monitor
The stage most often skipped, and the reason deployed models quietly rot.
Model performance degrades over time, always. Not because the model changes — it does not. Because the world does. Customer behaviour shifts, fraudsters adapt, products change, language changes. The model is a photograph of a moment that has passed.
Two kinds of drift, worth distinguishing:
Data drift — the inputs start looking different. Your customer base shifts younger; the model still works as designed but is now being asked about people it has not seen.
Concept drift — the relationship itself changes. What indicated fraud last year no longer does, because fraudsters read the same research you did.
What to watch: the statistical distribution of incoming data compared to training; prediction distributions (a sudden shift in how often the model says "yes" is a strong signal); accuracy wherever you can get real outcomes; and plain operational health — latency, errors, cost.
Getting true outcomes is often the hard part. You learn whether a loan defaulted two years later. Where ground truth is delayed or absent, monitoring the inputs for drift is your early warning system.
Why the loop matters
The arrow back from monitoring to data collection is the difference between a project and a product. Real systems are retrained on fresh data on a schedule — or when monitoring says drift has crossed a threshold.
This has an organisational consequence worth understanding before you start: an AI system is a commitment, not a delivery. Someone has to own it, watch it, and retrain it for as long as it runs. Teams that plan for a launch and not for a lifetime end up with a model nobody trusts and nobody dares turn off.
Check your understanding
0 of 3 answered
1.A model scores 99% on its training data but only 71% on the validation set. What is the most likely diagnosis?
2.You are predicting next week's sales from three years of daily data. How should you split it into training and test sets?
3.A new ticket-routing model is 87% accurate. What should you check before calling it a success?