Machine Learning Essentials

The Machine Learning Workflow


A data scientist spends three weeks on a churn model. The notebook reports 94% accuracy. The slide deck says 94%. The model ships. Six weeks later the retention team reports that of the customers it flagged, roughly six in ten actually churned — and they could have got close to that by phoning everyone whose contract expired that month.

Nothing was faked. The 94% was real, computed correctly, on data the model had never seen. And it was still meaningless, because of one line written on day four: the analyst filled in missing income values with the average income of the whole dataset, and only afterwards split the data into training and test sets. The test rows had quietly contributed their own averages to the numbers used to fill them in. Information flowed backwards. The test set stopped being a fair exam.

That is a workflow bug, not a modelling bug. No algorithm choice would have caught it, no hyperparameter search would have fixed it, and the model's internal metrics reported nothing wrong. This is the pattern behind most machine learning projects that fail in production: the steps were all performed, but in the wrong order, or with the wrong thing held back.

So the workflow is not bureaucracy. It is a specific sequence designed so that certain mistakes become impossible.

Where the split sits, and why it sits thereFrame theproblemLook at the dataSplit, thenstop lookingFitpreprocessingon train onlyBaseline, thentune with CVTouch thetest set onceEverything after the split is fitted on training rows; the test set is opened at the end.
The 94% notebook figure and the 60% field result differ only in where the split was placed.

The shape of the whole thing

Text
  1. Frame the problem        -> what decision changes because of this?  2. Collect and inspect      -> what do I actually have?  3. SPLIT                    <- everything before this line is allowed to see all data  4. Clean, engineer, encode  -> fitted on training data only  5. Baseline                 -> the number any model must beat  6. Train and tune           -> using cross-validation, never the test set  7. Final evaluation         -> test set, once  8. Deploy and monitor       -> the data will drift

Steps 4 through 6 loop many times. Step 7 happens once. The arrow at step 3 is the important one: it marks the moment after which the test set becomes untouchable.

Step 1: Frame the problem before you touch data

The question is never "can I build a model on this data?". It is "what decision will be made differently, by whom, and what does being wrong cost?".

Consider a hospital that wants to "predict patient readmission". That is not yet a specification. Push on it:

  • Who acts on the output? A discharge nurse, deciding whether to schedule a follow-up call.
  • When? At the moment of discharge — so any feature recorded after discharge is unusable, however predictive.
  • How many can they act on? The team can make 40 calls a day. So the useful output is not "will this patient be readmitted" but "rank today's discharges and flag the top 40".
  • What does each error cost? A missed readmission is a patient back in A&E. A false alarm is a five-minute phone call. Those costs are wildly asymmetric, which means overall accuracy is the wrong thing to optimise.

Notice how much of the technical design that conversation just settled. The prediction time fixed which features are legal. The capacity constraint turned a classification problem into a ranking problem. The cost asymmetry ruled out accuracy as a target metric.

Two metrics exist in every project: the business metric (readmissions prevented) and the model metric (recall at the top 40). Your job is to choose a model metric that moves the business metric. Teams that skip this optimise a number nobody asked for.

The leakage question, asked early

For every candidate feature, ask: would this value be known, and populated, at the exact moment the prediction is needed?

A column called discharge_summary_length is written after discharge. A column called total_visits_this_year includes the readmission you are trying to predict. A column called assigned_case_manager is only filled in for patients someone already flagged as high-risk — so it encodes the answer via a human's prior judgement.

Each of these will make your validation score go up and your production performance go down. Catching them costs ten minutes at the start and is nearly impossible to detect later, because a leaking model looks like a brilliant model.

Step 2: Look at the data before modelling it

Exploratory analysis is not a ritual. It answers questions that change what you build.

Python
import pandas as pddf = pd.read_csv("patients.csv")print(df.shape)                      # how many rows, how many columnsprint(df.dtypes)                     # is 'age' a string because of a stray '?'print(df.isnull().mean().sort_values(ascending=False).head(10))print(df["readmitted"].value_counts(normalize=True))   # how imbalancedprint(df.describe())                 # min/max sanity: negative ages, ages of 300print(df.duplicated().sum())         # duplicated rows inflate every score

Things that regularly turn up and change the plan:

What you findWhat it forces
Target is 3% positiveAccuracy is useless; use precision/recall, consider class weighting
A column is 80% missingUsually drop it; "missing" itself may be a useful flag
Ages of 0 and 999Sentinel values encoding "unknown" — must be converted to NaN, not averaged in
Two features correlate at 0.98Redundant; keep one, or coefficients become uninterpretable
A feature perfectly predicts the targetAlmost certainly leakage. Investigate before celebrating
Rows have timestampsYou must split by time, not randomly
Multiple rows per patientRandom splitting puts the same patient in train and test — split by patient

That last one is subtle and expensive. If a patient appears in ten rows and a random split scatters them across train and test, your model can memorise that specific patient rather than learning a general pattern. The fix is to split on the entity, not the row.

Step 3: Split, and then stop looking

The test set exists to answer one question: how will this perform on data it has never encountered? That answer is only honest if the data was genuinely never encountered — not by the model, and not by you.

Every decision you make after seeing test performance leaks a little information from the test set into the model. Try twelve architectures, pick the one with the best test score, and that score is no longer an estimate of future performance. It is the maximum of twelve noisy numbers, which is biased upward by construction.

Hence three sets, not two:

SetTypical sizeUsed forHow often
Training60–80%Fitting model parametersConstantly
Validation10–20%Choosing models, tuning hyperparametersMany times
Test10–20%Final honest estimateOnce, at the end

Which kind of split

  • Random — the default, valid when rows are independent.
  • Stratified — preserves the class balance in each split. Essential with rare targets: a random split of a 2%-positive dataset can easily produce a test set with a different positive rate, making the score noise.
  • Grouped — all rows for one patient, one customer, one device stay together.
  • Time-based — train on the past, test on the future. Any problem where you predict forward in time must use this. A random split lets the model train on Tuesday to predict Monday, which it will never be able to do in production.
Python
from sklearn.model_selection import train_test_splitX = df.drop(columns=["readmitted"])y = df["readmitted"]# hold out the test set first, and stratify because positives are rareX_temp, X_test, y_temp, y_test = train_test_split(    X, y, test_size=0.2, stratify=y, random_state=42)# then carve a validation set out of what remainsX_train, X_val, y_train, y_val = train_test_split(    X_temp, y_temp, test_size=0.25, stratify=y_temp, random_state=42)# 60 / 20 / 20

Step 4: Preprocessing, fitted on training data only

This is where the opening story went wrong, so it is worth being precise about the rule.

Preprocessing steps have parameters learned from data: the mean used to fill missing values, the mean and standard deviation used to scale, the vocabulary used to one-hot encode, the quantile boundaries used to bin. Every one of those must be computed from the training set and then applied unchanged to validation and test.

Concretely, with a training mean income of £41,200 and a test set that happens to contain three people earning £1 million each:

WrongRight
Compute mean onAll 10,000 rows → £41,4908,000 training rows → £41,200
Fill training gaps with£41,490£41,200
Fill test gaps with£41,490£41,200
Test set influenced its own inputs?YesNo
Reported scoreOptimistic by an unknown amountHonest

In scikit-learn the distinction is exactly the difference between fit_transform and transform. fit_transform on the training set; transform — never fit — on everything else.

The reliable way to never get this wrong is to stop doing it by hand and put every step inside a Pipeline. A pipeline is a single object containing the preprocessing and the model; calling fit on it fits every stage on training data only, and calling predict applies every stage in the same order. It becomes structurally impossible to leak.

Python
from sklearn.pipeline import Pipelinefrom sklearn.compose import ColumnTransformerfrom sklearn.impute import SimpleImputerfrom sklearn.preprocessing import StandardScaler, OneHotEncoderfrom sklearn.ensemble import RandomForestClassifiernumeric = ["age", "days_in_hospital", "num_medications"]categorical = ["admission_type", "discharge_disposition"]pre = ColumnTransformer([    ("num", Pipeline([        ("impute", SimpleImputer(strategy="median")),        ("scale", StandardScaler()),    ]), numeric),    ("cat", Pipeline([        ("impute", SimpleImputer(strategy="most_frequent")),        ("encode", OneHotEncoder(handle_unknown="ignore")),    ]), categorical),])model = Pipeline([    ("pre", pre),    ("clf", RandomForestClassifier(n_estimators=300, class_weight="balanced",                                   random_state=42)),])model.fit(X_train, y_train)     # every transformer fits on training data alone

Note handle_unknown="ignore". Without it, a category that appears in production but not in training crashes the encoder. That is a Tuesday-afternoon outage waiting to happen.

Feature engineering

The features you construct usually matter more than the algorithm you pick. Ratios, differences, counts, and time-since values are where most gains live: debt_to_income rather than debt and income separately; days_since_last_admission rather than a raw date; medications_per_day rather than totals that scale with length of stay.

The same discipline applies: any engineered feature computed using statistics across rows (a group mean, a target encoding) must be computed from training rows only.

Step 5: Establish a baseline before touching a real model

A baseline is the dumbest defensible prediction. It exists to give your model something to beat and to catch nonsense early.

  • Classification: always predict the majority class. On a dataset that is 89% negative, this scores 89% accuracy — which instantly tells you that a model reporting 90% accuracy has learned almost nothing.
  • Regression: always predict the training mean. This gives an R² of exactly 0 on the training data, and about 0 (often slightly below) on new data.
  • Better baseline: the existing rule the business already uses. If nurses currently flag anyone over 75 with more than two prior admissions, that rule is your real competitor.
Python
from sklearn.dummy import DummyClassifierfrom sklearn.metrics import recall_scoredummy = DummyClassifier(strategy="most_frequent").fit(X_train, y_train)print("baseline accuracy:", dummy.score(X_val, y_val))print("baseline recall:", recall_score(y_val, dummy.predict(X_val)))  # 0.0

That recall of zero is the point. The baseline is 89% accurate and catches not a single readmission — which is exactly why accuracy was the wrong metric for this problem, and now you have the evidence rather than the assertion.

Step 6: Train, compare, tune — with cross-validation

A single validation set gives one noisy estimate. Cross-validation gives several: split the training data into k folds, train on k−1 and validate on the held-out one, rotate, and average. With k=5 you get five estimates and, importantly, their spread.

Python
from sklearn.model_selection import cross_val_score, StratifiedKFoldfrom sklearn.linear_model import LogisticRegressioncv = StratifiedKFold(n_splits=5, shuffle=True, random_state=42)for name, clf in [("logreg", LogisticRegression(max_iter=1000, class_weight="balanced")),                  ("forest", RandomForestClassifier(n_estimators=300,                                                    class_weight="balanced",                                                    random_state=42))]:    pipe = Pipeline([("pre", pre), ("clf", clf)])    scores = cross_val_score(pipe, X_train, y_train, cv=cv, scoring="recall")    print(f"{name}: {scores.mean():.3f} +/- {scores.std():.3f}")

Read both numbers. A model scoring 0.72 ± 0.02 is not worse than one scoring 0.74 ± 0.09: the second one's advantage is inside its own noise, and the first is far more predictable. Reporting only the mean is how teams convince themselves of improvements that do not exist.

Hyperparameter tuning then searches for settings, still using cross-validation on the training data:

Python
from sklearn.model_selection import RandomizedSearchCVparams = {    "clf__n_estimators": [200, 400, 800],    "clf__max_depth": [None, 8, 16, 24],    "clf__min_samples_leaf": [1, 5, 20],}search = RandomizedSearchCV(model, params, n_iter=20, cv=cv,                            scoring="recall", n_jobs=-1, random_state=42)search.fit(X_train, y_train)print(search.best_params_, search.best_score_)

Randomised search beats exhaustive grid search in most practical settings: with a fixed compute budget it explores more distinct values of the parameters that actually matter, instead of exhaustively varying ones that do not.

Step 7: The test set, once

When you have chosen a model and its settings, refit on all the training-plus-validation data and evaluate on the test set exactly once.

Python
from sklearn.metrics import classification_report, confusion_matrixfinal = search.best_estimator_.fit(X_temp, y_temp)   # train + validationy_pred = final.predict(X_test)print(confusion_matrix(y_test, y_pred))print(classification_report(y_test, y_pred, digits=3))

If that number disappoints you, the honest move is to accept it, or to go back and improve the model using cross-validation on the training data — and then acknowledge that your test estimate is now slightly optimistic. The dishonest move, which is extremely common, is to tweak until the test number looks good and report it as though it were untouched.

Every time you look at the test set and change something, it becomes a little more like a validation set and a little less like a fair estimate of the future.

Step 8: Deployment is when the clock starts

A trained model is a file. Shipping it means saving the entire pipeline — preprocessing included — and serving it behind something that receives raw inputs in the same form the training data had.

Python
import joblibjoblib.dump(final, "readmission_pipeline.joblib")   # preprocessing travels with it

Saving only the classifier and re-implementing the preprocessing in the serving code is a reliable way to produce a model that works in the notebook and returns nonsense in production, because the two implementations drift apart within weeks.

Then monitor. Models decay, not because the code rots but because the world moves:

What driftsExampleDetect by
Input distributionAverage patient age rises after a new ward opensComparing live feature distributions to training ones
Relationship to targetA new discharge protocol changes what causes readmissionTracking live accuracy once outcomes arrive
Upstream schemaA field starts arriving in minutes instead of hoursRange and type checks on every request
Feedback loopsFlagged patients get calls, so they stop being readmittedHolding out a small untreated control group

That last row is worth pausing on. A successful intervention destroys the very pattern the model learned. If flagged patients receive follow-up calls and consequently do not return, the model's predictions start looking wrong — precisely because they were right and acted upon. Without a control group you cannot tell that apart from a broken model.

Where the time actually goes

StageShare of project timeWhere it goes wrong
Framing5–10%Skipped; team optimises the wrong metric
Data collection and cleaning50–70%Underestimated in every plan ever written
Feature engineering10–20%Leakage introduced here
Model training and tuning5–15%Over-invested; the fun part
Evaluation5%Test set peeked at repeatedly
Deployment and monitoring10–20%Treated as someone else's job

Beginners routinely invert the middle two rows, spending days on hyperparameters and an hour on data quality. The returns are the other way round almost every time.

What this means when you build something

Set the project up so the expensive mistakes cannot happen, rather than trying to remember not to make them.

Split the data in the first thirty minutes, before you have formed any opinion about it, and write the test set to a separate file you do not open. Put every transformation inside a pipeline so that fitting on the wrong rows is not something you can do by accident. Compute the dumb baseline before the clever model, so you know what "good" means. Decide the metric in the framing conversation, with the person who will act on the output, and write it down where you will see it.

Then, when your model reports 94%, you will be able to say precisely what that number means — and whether anyone should believe it.