Course Content
Machine Learning Essentials
6 sections · 16 lessons
The Machine Learning Workflow
A data scientist spends three weeks on a churn model. The notebook reports 94% accuracy. The slide deck says 94%. The model ships. Six weeks later the retention team reports that of the customers it flagged, roughly six in ten actually churned — and they could have got close to that by phoning everyone whose contract expired that month.
Nothing was faked. The 94% was real, computed correctly, on data the model had never seen. And it was still meaningless, because of one line written on day four: the analyst filled in missing income values with the average income of the whole dataset, and only afterwards split the data into training and test sets. The test rows had quietly contributed their own averages to the numbers used to fill them in. Information flowed backwards. The test set stopped being a fair exam.
That is a workflow bug, not a modelling bug. No algorithm choice would have caught it, no hyperparameter search would have fixed it, and the model's internal metrics reported nothing wrong. This is the pattern behind most machine learning projects that fail in production: the steps were all performed, but in the wrong order, or with the wrong thing held back.
So the workflow is not bureaucracy. It is a specific sequence designed so that certain mistakes become impossible.
The shape of the whole thing
1. Frame the problem -> what decision changes because of this? 2. Collect and inspect -> what do I actually have? 3. SPLIT <- everything before this line is allowed to see all data 4. Clean, engineer, encode -> fitted on training data only 5. Baseline -> the number any model must beat 6. Train and tune -> using cross-validation, never the test set 7. Final evaluation -> test set, once 8. Deploy and monitor -> the data will driftSteps 4 through 6 loop many times. Step 7 happens once. The arrow at step 3 is the important one: it marks the moment after which the test set becomes untouchable.
Step 1: Frame the problem before you touch data
The question is never "can I build a model on this data?". It is "what decision will be made differently, by whom, and what does being wrong cost?".
Consider a hospital that wants to "predict patient readmission". That is not yet a specification. Push on it:
- Who acts on the output? A discharge nurse, deciding whether to schedule a follow-up call.
- When? At the moment of discharge — so any feature recorded after discharge is unusable, however predictive.
- How many can they act on? The team can make 40 calls a day. So the useful output is not "will this patient be readmitted" but "rank today's discharges and flag the top 40".
- What does each error cost? A missed readmission is a patient back in A&E. A false alarm is a five-minute phone call. Those costs are wildly asymmetric, which means overall accuracy is the wrong thing to optimise.
Notice how much of the technical design that conversation just settled. The prediction time fixed which features are legal. The capacity constraint turned a classification problem into a ranking problem. The cost asymmetry ruled out accuracy as a target metric.
Two metrics exist in every project: the business metric (readmissions prevented) and the model metric (recall at the top 40). Your job is to choose a model metric that moves the business metric. Teams that skip this optimise a number nobody asked for.
The leakage question, asked early
For every candidate feature, ask: would this value be known, and populated, at the exact moment the prediction is needed?
A column called discharge_summary_length is written after discharge. A column called total_visits_this_year includes the readmission you are trying to predict. A column called assigned_case_manager is only filled in for patients someone already flagged as high-risk — so it encodes the answer via a human's prior judgement.
Each of these will make your validation score go up and your production performance go down. Catching them costs ten minutes at the start and is nearly impossible to detect later, because a leaking model looks like a brilliant model.
Step 2: Look at the data before modelling it
Exploratory analysis is not a ritual. It answers questions that change what you build.
1import pandas as pd23df = pd.read_csv("patients.csv")45print(df.shape) # how many rows, how many columns6print(df.dtypes) # is 'age' a string because of a stray '?'7print(df.isnull().mean().sort_values(ascending=False).head(10))8print(df["readmitted"].value_counts(normalize=True)) # how imbalanced9print(df.describe()) # min/max sanity: negative ages, ages of 30010print(df.duplicated().sum()) # duplicated rows inflate every scoreThings that regularly turn up and change the plan:
| What you find | What it forces |
|---|---|
| Target is 3% positive | Accuracy is useless; use precision/recall, consider class weighting |
| A column is 80% missing | Usually drop it; "missing" itself may be a useful flag |
| Ages of 0 and 999 | Sentinel values encoding "unknown" — must be converted to NaN, not averaged in |
| Two features correlate at 0.98 | Redundant; keep one, or coefficients become uninterpretable |
| A feature perfectly predicts the target | Almost certainly leakage. Investigate before celebrating |
| Rows have timestamps | You must split by time, not randomly |
| Multiple rows per patient | Random splitting puts the same patient in train and test — split by patient |
That last one is subtle and expensive. If a patient appears in ten rows and a random split scatters them across train and test, your model can memorise that specific patient rather than learning a general pattern. The fix is to split on the entity, not the row.
Step 3: Split, and then stop looking
The test set exists to answer one question: how will this perform on data it has never encountered? That answer is only honest if the data was genuinely never encountered — not by the model, and not by you.
Every decision you make after seeing test performance leaks a little information from the test set into the model. Try twelve architectures, pick the one with the best test score, and that score is no longer an estimate of future performance. It is the maximum of twelve noisy numbers, which is biased upward by construction.
Hence three sets, not two:
| Set | Typical size | Used for | How often |
|---|---|---|---|
| Training | 60–80% | Fitting model parameters | Constantly |
| Validation | 10–20% | Choosing models, tuning hyperparameters | Many times |
| Test | 10–20% | Final honest estimate | Once, at the end |
Which kind of split
- Random — the default, valid when rows are independent.
- Stratified — preserves the class balance in each split. Essential with rare targets: a random split of a 2%-positive dataset can easily produce a test set with a different positive rate, making the score noise.
- Grouped — all rows for one patient, one customer, one device stay together.
- Time-based — train on the past, test on the future. Any problem where you predict forward in time must use this. A random split lets the model train on Tuesday to predict Monday, which it will never be able to do in production.
1from sklearn.model_selection import train_test_split23X = df.drop(columns=["readmitted"])4y = df["readmitted"]56# hold out the test set first, and stratify because positives are rare7X_temp, X_test, y_temp, y_test = train_test_split(8 X, y, test_size=0.2, stratify=y, random_state=429)10# then carve a validation set out of what remains11X_train, X_val, y_train, y_val = train_test_split(12 X_temp, y_temp, test_size=0.25, stratify=y_temp, random_state=4213)14# 60 / 20 / 20Step 4: Preprocessing, fitted on training data only
This is where the opening story went wrong, so it is worth being precise about the rule.
Preprocessing steps have parameters learned from data: the mean used to fill missing values, the mean and standard deviation used to scale, the vocabulary used to one-hot encode, the quantile boundaries used to bin. Every one of those must be computed from the training set and then applied unchanged to validation and test.
Concretely, with a training mean income of £41,200 and a test set that happens to contain three people earning £1 million each:
| Wrong | Right | |
|---|---|---|
| Compute mean on | All 10,000 rows → £41,490 | 8,000 training rows → £41,200 |
| Fill training gaps with | £41,490 | £41,200 |
| Fill test gaps with | £41,490 | £41,200 |
| Test set influenced its own inputs? | Yes | No |
| Reported score | Optimistic by an unknown amount | Honest |
In scikit-learn the distinction is exactly the difference between fit_transform and transform. fit_transform on the training set; transform — never fit — on everything else.
The reliable way to never get this wrong is to stop doing it by hand and put every step inside a Pipeline. A pipeline is a single object containing the preprocessing and the model; calling fit on it fits every stage on training data only, and calling predict applies every stage in the same order. It becomes structurally impossible to leak.
1from sklearn.pipeline import Pipeline2from sklearn.compose import ColumnTransformer3from sklearn.impute import SimpleImputer4from sklearn.preprocessing import StandardScaler, OneHotEncoder5from sklearn.ensemble import RandomForestClassifier67numeric = ["age", "days_in_hospital", "num_medications"]8categorical = ["admission_type", "discharge_disposition"]910pre = ColumnTransformer([11 ("num", Pipeline([12 ("impute", SimpleImputer(strategy="median")),13 ("scale", StandardScaler()),14 ]), numeric),15 ("cat", Pipeline([16 ("impute", SimpleImputer(strategy="most_frequent")),17 ("encode", OneHotEncoder(handle_unknown="ignore")),18 ]), categorical),19])2021model = Pipeline([22 ("pre", pre),23 ("clf", RandomForestClassifier(n_estimators=300, class_weight="balanced",24 random_state=42)),25])2627model.fit(X_train, y_train) # every transformer fits on training data aloneNote handle_unknown="ignore". Without it, a category that appears in production but not in training crashes the encoder. That is a Tuesday-afternoon outage waiting to happen.
Feature engineering
The features you construct usually matter more than the algorithm you pick. Ratios, differences, counts, and time-since values are where most gains live: debt_to_income rather than debt and income separately; days_since_last_admission rather than a raw date; medications_per_day rather than totals that scale with length of stay.
The same discipline applies: any engineered feature computed using statistics across rows (a group mean, a target encoding) must be computed from training rows only.
Step 5: Establish a baseline before touching a real model
A baseline is the dumbest defensible prediction. It exists to give your model something to beat and to catch nonsense early.
- Classification: always predict the majority class. On a dataset that is 89% negative, this scores 89% accuracy — which instantly tells you that a model reporting 90% accuracy has learned almost nothing.
- Regression: always predict the training mean. This gives an R² of exactly 0 on the training data, and about 0 (often slightly below) on new data.
- Better baseline: the existing rule the business already uses. If nurses currently flag anyone over 75 with more than two prior admissions, that rule is your real competitor.
1from sklearn.dummy import DummyClassifier2from sklearn.metrics import recall_score34dummy = DummyClassifier(strategy="most_frequent").fit(X_train, y_train)5print("baseline accuracy:", dummy.score(X_val, y_val))6print("baseline recall:", recall_score(y_val, dummy.predict(X_val))) # 0.0That recall of zero is the point. The baseline is 89% accurate and catches not a single readmission — which is exactly why accuracy was the wrong metric for this problem, and now you have the evidence rather than the assertion.
Step 6: Train, compare, tune — with cross-validation
A single validation set gives one noisy estimate. Cross-validation gives several: split the training data into k folds, train on k−1 and validate on the held-out one, rotate, and average. With k=5 you get five estimates and, importantly, their spread.
1from sklearn.model_selection import cross_val_score, StratifiedKFold2from sklearn.linear_model import LogisticRegression34cv = StratifiedKFold(n_splits=5, shuffle=True, random_state=42)56for name, clf in [("logreg", LogisticRegression(max_iter=1000, class_weight="balanced")),7 ("forest", RandomForestClassifier(n_estimators=300,8 class_weight="balanced",9 random_state=42))]:10 pipe = Pipeline([("pre", pre), ("clf", clf)])11 scores = cross_val_score(pipe, X_train, y_train, cv=cv, scoring="recall")12 print(f"{name}: {scores.mean():.3f} +/- {scores.std():.3f}")Read both numbers. A model scoring 0.72 ± 0.02 is not worse than one scoring 0.74 ± 0.09: the second one's advantage is inside its own noise, and the first is far more predictable. Reporting only the mean is how teams convince themselves of improvements that do not exist.
Hyperparameter tuning then searches for settings, still using cross-validation on the training data:
1from sklearn.model_selection import RandomizedSearchCV23params = {4 "clf__n_estimators": [200, 400, 800],5 "clf__max_depth": [None, 8, 16, 24],6 "clf__min_samples_leaf": [1, 5, 20],7}89search = RandomizedSearchCV(model, params, n_iter=20, cv=cv,10 scoring="recall", n_jobs=-1, random_state=42)11search.fit(X_train, y_train)12print(search.best_params_, search.best_score_)Randomised search beats exhaustive grid search in most practical settings: with a fixed compute budget it explores more distinct values of the parameters that actually matter, instead of exhaustively varying ones that do not.
Step 7: The test set, once
When you have chosen a model and its settings, refit on all the training-plus-validation data and evaluate on the test set exactly once.
1from sklearn.metrics import classification_report, confusion_matrix23final = search.best_estimator_.fit(X_temp, y_temp) # train + validation4y_pred = final.predict(X_test)56print(confusion_matrix(y_test, y_pred))7print(classification_report(y_test, y_pred, digits=3))If that number disappoints you, the honest move is to accept it, or to go back and improve the model using cross-validation on the training data — and then acknowledge that your test estimate is now slightly optimistic. The dishonest move, which is extremely common, is to tweak until the test number looks good and report it as though it were untouched.
Every time you look at the test set and change something, it becomes a little more like a validation set and a little less like a fair estimate of the future.
Step 8: Deployment is when the clock starts
A trained model is a file. Shipping it means saving the entire pipeline — preprocessing included — and serving it behind something that receives raw inputs in the same form the training data had.
import joblibjoblib.dump(final, "readmission_pipeline.joblib") # preprocessing travels with itSaving only the classifier and re-implementing the preprocessing in the serving code is a reliable way to produce a model that works in the notebook and returns nonsense in production, because the two implementations drift apart within weeks.
Then monitor. Models decay, not because the code rots but because the world moves:
| What drifts | Example | Detect by |
|---|---|---|
| Input distribution | Average patient age rises after a new ward opens | Comparing live feature distributions to training ones |
| Relationship to target | A new discharge protocol changes what causes readmission | Tracking live accuracy once outcomes arrive |
| Upstream schema | A field starts arriving in minutes instead of hours | Range and type checks on every request |
| Feedback loops | Flagged patients get calls, so they stop being readmitted | Holding out a small untreated control group |
That last row is worth pausing on. A successful intervention destroys the very pattern the model learned. If flagged patients receive follow-up calls and consequently do not return, the model's predictions start looking wrong — precisely because they were right and acted upon. Without a control group you cannot tell that apart from a broken model.
Where the time actually goes
| Stage | Share of project time | Where it goes wrong |
|---|---|---|
| Framing | 5–10% | Skipped; team optimises the wrong metric |
| Data collection and cleaning | 50–70% | Underestimated in every plan ever written |
| Feature engineering | 10–20% | Leakage introduced here |
| Model training and tuning | 5–15% | Over-invested; the fun part |
| Evaluation | 5% | Test set peeked at repeatedly |
| Deployment and monitoring | 10–20% | Treated as someone else's job |
Beginners routinely invert the middle two rows, spending days on hyperparameters and an hour on data quality. The returns are the other way round almost every time.
What this means when you build something
Set the project up so the expensive mistakes cannot happen, rather than trying to remember not to make them.
Split the data in the first thirty minutes, before you have formed any opinion about it, and write the test set to a separate file you do not open. Put every transformation inside a pipeline so that fitting on the wrong rows is not something you can do by accident. Compute the dumb baseline before the clever model, so you know what "good" means. Decide the metric in the framing conversation, with the person who will act on the output, and write it down where you will see it.
Then, when your model reports 94%, you will be able to say precisely what that number means — and whether anyone should believe it.