Course Content
Machine Learning Essentials
6 sections · 16 lessons
Capstone Project: End-to-End Machine Learning System
Someone opens your project repository. They have about four minutes before deciding whether to read further.
In the first version they see Untitled17.ipynb, 340 cells, several of which error, and a final cell printing 0.94. They cannot tell what was predicted, on what data, whether that number came from a test set, or what would happen if they wanted to use it. They close the tab.
In the second version they see a README stating the problem in two sentences, the headline metric with a comparison to a baseline, and an honest paragraph on limitations. Below it: a src/ directory, a saved pipeline, a Dockerfile, and a curl command that returns a prediction. They run the command. It works.
Both projects contain the same model. Only one demonstrates that you can build a system, and building a system is the thing being assessed here — end to end, from an ambiguous business question to something another person can call.
This piece of work asks you to do that once, properly, on a problem of your choosing.
What you will produce
| Artefact | What good looks like |
|---|---|
| Problem statement | Two paragraphs: the decision, who makes it, what an error costs |
| Data audit | Row counts, missingness, target balance, leakage checks performed |
| Training code | Runs start to finish from a clean checkout with one command |
| Saved pipeline | One file containing preprocessing and model together |
| Evaluation | Baseline, chosen metric, held-out score, per-segment breakdown |
| Error analysis | Twenty worst predictions inspected, with a written hypothesis |
| Serving layer | An API or app that loads the pipeline and returns predictions |
| Model card | Intended use, training data, metrics, limitations, ethical notes |
| README | Problem, result, how to run it, what you would do next |
Choosing a problem you will not regret
More capstones fail at the choice than at the modelling. Apply four tests before committing.
Can you state the decision? Complete this sentence: "When the model outputs X, someone does Y instead of Z." If you cannot, you have a dataset, not a problem.
Is the target real and available before the decision? Every feature must be knowable at prediction time. This kills more datasets than people expect.
Is there enough data? Roughly 1,000 rows minimum for a simple problem, and for classification at least 100 examples of the rarest class you care about. Fewer than that and your evaluation is noise.
Is it boring enough to finish? A tabular regression or binary classification you complete beats an ambitious multi-modal system you abandon at 60%.
| Problem | Type | Why it works | Watch out for |
|---|---|---|---|
| House price prediction | Regression | Rich features, intuitive, skewed target teaches log transforms | Outlier mansions dominating RMSE |
| Telecom churn | Classification | Clear business action, moderate imbalance | Features recorded after cancellation |
| Loan default | Classification | Asymmetric costs make threshold tuning meaningful | Fairness across demographic groups |
| Bike-share demand | Regression | Strong seasonal features, easy to engineer | Must split by time, never randomly |
| Hospital readmission | Classification | Heavy imbalance forces proper metrics | Leakage from discharge-time fields |
| Employee attrition | Classification | Small, clean, fast to iterate | Often too small for a reliable test set |
Avoid datasets where the winning approach is public and copy-pasteable. You will learn more from a messy dataset with an 0.72 AUC than from one where everybody scores 0.99.
Phase one: frame it, then audit the data
Write the problem statement before loading anything. Half a page:
- The decision being supported, and by whom.
- The prediction point — the exact moment the model is called.
- What a false positive costs. What a false negative costs. Numbers if you can get them, a ranking if you cannot.
- The metric that follows from those costs, and why not accuracy.
- What "good enough to be useful" would be.
That last item is the one people skip and later regret. Deciding in advance that anything above 0.75 recall at 20% precision is worth shipping stops you from moving the goalposts once you see your results.
Then audit. Not exploratory plotting for its own sake — specific questions with recorded answers.
1import pandas as pd23df = pd.read_csv("data/raw/dataset.csv")45print("shape:", df.shape)6print("duplicates:", df.duplicated().sum())7print("\nmissing (top 10):")8print((df.isna().mean() * 100).sort_values(ascending=False).head(10).round(1))9print("\ntarget balance:")10print(df[TARGET].value_counts(normalize=True).round(4))11print("\nsuspicious values:")12print(df.describe().T[["min", "max"]])1314# the leakage screen: any single feature that predicts the target too well15from sklearn.feature_selection import mutual_info_classif16mi = mutual_info_classif(df[NUMERIC].fillna(-1), df[TARGET])17print(pd.Series(mi, index=NUMERIC).sort_values(ascending=False).head(10))Record what you find in a short document. Two findings deserve explicit attention:
Sentinel values. Ages of 0 or 999, incomes of −1, dates of 1900-01-01. These encode "unknown" and must become NaN before any imputation, or you will average a 999 into your fill value.
Anything that predicts too well. A single feature with an AUC above 0.95 on its own is leakage until proven otherwise. Investigate before celebrating — this is the failure that makes a capstone worthless while looking spectacular.
Phase two: split first, then build the pipeline
Split before you clean anything, and choose the split type deliberately.
| Split | Use when | What goes wrong otherwise |
|---|---|---|
| Stratified random | Independent rows, imbalanced target | Test set gets a different class rate; scores become noise |
| Grouped | Multiple rows per customer, patient, device | The same entity appears in train and test; the model memorises it |
| Time-based | Anything predicting forward in time | Model trains on the future to predict the past |
1from sklearn.model_selection import train_test_split23X, y = df.drop(columns=[TARGET]), df[TARGET]45X_temp, X_test, y_temp, y_test = train_test_split(6 X, y, test_size=0.2, stratify=y, random_state=42)78X_test.assign(**{TARGET: y_test}).to_parquet("data/holdout.parquet")9# and do not open it again until the very endThen put every transformation inside a pipeline so that nothing is ever fitted on data it should not see.
1from sklearn.pipeline import Pipeline2from sklearn.compose import ColumnTransformer3from sklearn.impute import SimpleImputer4from sklearn.preprocessing import StandardScaler, OneHotEncoder56pre = ColumnTransformer([7 ("num", Pipeline([("impute", SimpleImputer(strategy="median")),8 ("scale", StandardScaler())]), NUMERIC),9 ("cat", Pipeline([("impute", SimpleImputer(strategy="most_frequent")),10 ("encode", OneHotEncoder(handle_unknown="ignore",11 min_frequency=10))]), CATEGORICAL),12])min_frequency=10 groups rare categories together, which stops a category seen twice in training from producing a column the model cannot learn anything about. handle_unknown="ignore" stops an unseen category crashing the service later.
Feature engineering is where the marks are
Algorithm choice is worth a few percent. Features are worth more. Spend real time here and document each one.
1def add_features(df: pd.DataFrame) -> pd.DataFrame:2 out = df.copy()3 out["debt_to_income"] = out["debt"] / out["income"].clip(lower=1)4 out["spend_per_product"] = out["annual_spend"] / out["num_products"]5 out["days_since_contact"] = (6 pd.Timestamp("2026-01-01") - pd.to_datetime(out["last_contact"])7 ).dt.days8 out["is_new"] = (out["tenure_years"] < 1).astype(int)9 return outTwo rules. Ratios and differences almost always beat their raw components, because they encode a relationship the model would otherwise have to discover. And any feature built from statistics across rows — a group mean, a target encoding — must be computed inside cross-validation folds, or it leaks.
Phase three: baseline first, then models
Compute the dumb baseline before anything else, so that every later number has a reference point.
1from sklearn.dummy import DummyClassifier2from sklearn.model_selection import cross_val_score, StratifiedKFold34cv = StratifiedKFold(5, shuffle=True, random_state=42)5dummy = DummyClassifier(strategy="prior")6print("baseline:", cross_val_score(dummy, X_temp, y_temp, cv=cv,7 scoring="average_precision").mean())Then work up in complexity, recording every result in a table you keep as you go.
1from sklearn.linear_model import LogisticRegression2from sklearn.ensemble import RandomForestClassifier, HistGradientBoostingClassifier34candidates = {5 "logistic": LogisticRegression(max_iter=2000, class_weight="balanced"),6 "random_forest": RandomForestClassifier(n_estimators=400, n_jobs=-1,7 class_weight="balanced_subsample",8 random_state=42),9 "hist_gbm": HistGradientBoostingClassifier(max_iter=400, random_state=42),10}1112rows = []13for name, clf in candidates.items():14 pipe = Pipeline([("pre", pre), ("clf", clf)])15 s = cross_val_score(pipe, X_temp, y_temp, cv=cv,16 scoring="average_precision", n_jobs=-1)17 rows.append({"model": name, "ap_mean": s.mean(), "ap_std": s.std()})1819print(pd.DataFrame(rows).sort_values("ap_mean", ascending=False))Report the standard deviation, not just the mean. A model at 0.61 ± 0.02 is not worse than one at 0.63 ± 0.08: the second one's lead is inside its own noise, and claiming it is the most common way capstones overstate their results.
Tune only the top one or two candidates, with randomised search over a log-spaced grid, and keep the tuning inside cross-validation on the training data.
Phase four: evaluate honestly
Now open the held-out file. Once.
1final = best_pipeline.fit(X_temp, y_temp)23holdout = pd.read_parquet("data/holdout.parquet")4X_hold, y_hold = holdout.drop(columns=[TARGET]), holdout[TARGET]5probs = final.predict_proba(X_hold)[:, 1]Three analyses turn a score into an argument.
Threshold selection against real costs
Choose the threshold without looking at the holdout results, using out-of-fold predictions on the training data — every row predicted by a model that never saw it.
1import numpy as np2from sklearn.model_selection import cross_val_predict34oof = cross_val_predict(best_pipeline, X_temp, y_temp, cv=cv,5 method="predict_proba")[:, 1]67COST_FN, COST_FP = 220.0, 4.08rows = []9for t in np.arange(0.02, 0.99, 0.01):10 pred = oof >= t11 fn = int(((y_temp == 1) & ~pred).sum())12 fp = int(((y_temp == 0) & pred).sum())13 rows.append({"threshold": round(t, 2), "fn": fn, "fp": fp,14 "cost": fn * COST_FN + fp * COST_FP})15costs = pd.DataFrame(rows)16print(costs.nsmallest(5, "cost"))Set CHOSEN_THRESHOLD from this table, then apply it unchanged to the holdout, and report it as part of the model. A model shipped at 0.5 because that is the default has an unexamined assumption that both errors cost the same.
Segment breakdown
1holdout["prob"] = probs2holdout["pred"] = (probs >= CHOSEN_THRESHOLD).astype(int)34for col in ["region", "age_band", "tenure_band"]:5 print(holdout.groupby(col).apply(6 lambda g: pd.Series({7 "n": len(g),8 "positive_rate": g[TARGET].mean(),9 "recall": (g[(g[TARGET] == 1)]["pred"]).mean(),10 "precision": g[g["pred"] == 1][TARGET].mean(),11 }), include_groups=False).round(3))An aggregate AUC of 0.87 can hide 0.91 on one region and 0.58 on another. If any segment corresponds to a protected characteristic, this table is not optional — it is the difference between a model and a liability.
Error analysis
Sort by absolute error, take the twenty worst, and look at the actual rows.
worst = holdout.assign(err=(holdout[TARGET] - holdout["prob"]).abs()) \ .nlargest(20, "err")print(worst[FEATURES + [TARGET, "prob"]])Then write a paragraph on what you see. Are they all from one region? All recent sign-ups? All missing the same field? This paragraph is frequently the most impressive thing in a capstone, because it shows you interrogated the model rather than just scoring it — and it usually suggests the next feature.
Phase five: package and serve
Save the whole pipeline, plus metadata, plus a fixture for verification.
1import joblib, json, sklearn, sys2from datetime import datetime, timezone34joblib.dump(final, "artifacts/model_v1.0.0.joblib")56json.dump({7 "version": "1.0.0",8 "trained_at": datetime.now(timezone.utc).isoformat(),9 "sklearn": sklearn.__version__,10 "python": sys.version.split()[0],11 "features": list(X_temp.columns),12 "decision_threshold": CHOSEN_THRESHOLD,13 "holdout_metrics": {"average_precision": 0.612, "roc_auc": 0.871},14 "baseline_average_precision": 0.083,15}, open("artifacts/model_v1.0.0.meta.json", "w"), indent=2)1617fixture = X_hold.head(5)18json.dump({"inputs": fixture.to_dict("records"),19 "expected": final.predict_proba(fixture)[:, 1].round(6).tolist()},20 open("artifacts/fixture.json", "w"), indent=2)That fixture is checked at service startup. If the loaded model produces different numbers — because someone upgraded a library or deployed the wrong file — the service fails loudly instead of quietly returning wrong answers for a month.
Then the service itself. Keep it small; it is judged on correctness, not features.
1from fastapi import FastAPI2from pydantic import BaseModel, Field3from typing import Literal4import joblib, pandas as pd56app = FastAPI(title="Churn API", version="1.0.0")7MODEL = joblib.load("artifacts/model_v1.0.0.joblib")8THRESHOLD = 0.34910class Features(BaseModel):11 income: float = Field(ge=0, le=1_000_000)12 age: int = Field(ge=18, le=100)13 tenure_years: float = Field(ge=0, le=60)14 num_products: int = Field(ge=1, le=10)15 region: Literal["north", "south", "east", "west"]1617@app.get("/health")18def health() -> dict:19 return {"status": "ok", "model_version": "1.0.0"}2021@app.post("/predict")22def predict(f: Features) -> dict:23 p = float(MODEL.predict_proba(pd.DataFrame([f.model_dump()]))[0, 1])24 return {"probability": round(p, 4), "will_churn": p >= THRESHOLD,25 "threshold": THRESHOLD, "model_version": "1.0.0"}Add a Dockerfile with pinned dependencies and confirm the container starts and answers a request. The README should contain the exact docker build, docker run, and curl commands, and they should work when pasted.
Phase six: write it down
The model card is a page and takes forty minutes.
| Section | Contents |
|---|---|
| Intended use | The decision it supports; who uses it; how often |
| Out of scope | Uses it must not be put to — credit decisions, hiring, anything legal |
| Training data | Source, date range, row count, known biases in collection |
| Features | List, with anything ethically sensitive flagged |
| Performance | Metric, held-out value, baseline, per-segment table |
| Limitations | Where it fails; the extrapolation range; what would degrade it |
| Monitoring | What to track and what should trigger retraining |
Write the limitations section as if a sceptical colleague will read it, because stating a weakness plainly reads as competence while omitting it reads as not having looked.
The mistakes that sink these projects
| Mistake | How it shows up | Prevention |
|---|---|---|
| Leakage | AUC of 0.99 that nobody questions | Ask of every feature: known before the decision? |
| Test set used repeatedly | Reported score is the best of many attempts | Write the holdout to a file and leave it shut |
| Preprocessing outside the pipeline | Serving code re-implements scaling and drifts | Everything in Pipeline, saved as one object |
| Accuracy on imbalanced data | 96% accuracy, zero positives caught | Average precision or balanced accuracy |
| No baseline | No way to tell whether the model helps | DummyClassifier in the first hour |
| Tuning before features | Days of grid search for 0.4% | Feature work first; tune once, at the end |
| Random split on time-series data | Excellent scores, useless model | Split by date |
| Only aggregate metrics | Fails badly on a segment nobody checked | Per-segment table |
| Nothing runs from a clean checkout | Reviewer cannot reproduce anything | Test it in a fresh clone before submitting |
A schedule that finishes
| Stage | Share of effort | Done when |
|---|---|---|
| Framing and data audit | 15% | Problem statement written; leakage screen run |
| Split and pipeline | 10% | Holdout saved; pipeline transforms without error |
| Features and models | 30% | Comparison table with means and standard deviations |
| Evaluation and error analysis | 20% | Threshold chosen; segments checked; 20 errors read |
| Packaging and serving | 15% | curl against the container returns a prediction |
| Writing up | 10% | README and model card complete |
The temptation is to spend 60% on modelling. Resist it. The modelling is the part most easily done well with least effort, and the parts people skip — framing, error analysis, packaging, writing — are exactly the parts that distinguish a finished system from a notebook.
What this means when you build something
Work so that every claim you make is checkable. That means the holdout file stays closed until the end, the baseline is computed first so the headline number has a reference, the pipeline is one saved object so nobody can ask whether the serving code matches training, and the fixture test proves the deployed artefact behaves like the trained one.
Then be honest in the write-up. A project reporting 0.61 average precision against a 0.08 baseline, with a clear account of where it fails and which segment it underserves, is stronger than one reporting 0.99 with no baseline and no error analysis — because the first is a system somebody could deploy, and the second is a number nobody can trust.