Machine Learning Essentials

Capstone Project: End-to-End Machine Learning System


Someone opens your project repository. They have about four minutes before deciding whether to read further.

In the first version they see Untitled17.ipynb, 340 cells, several of which error, and a final cell printing 0.94. They cannot tell what was predicted, on what data, whether that number came from a test set, or what would happen if they wanted to use it. They close the tab.

In the second version they see a README stating the problem in two sentences, the headline metric with a comparison to a baseline, and an honest paragraph on limitations. Below it: a src/ directory, a saved pipeline, a Dockerfile, and a curl command that returns a prediction. They run the command. It works.

Both projects contain the same model. Only one demonstrates that you can build a system, and building a system is the thing being assessed here — end to end, from an ambiguous business question to something another person can call.

This piece of work asks you to do that once, properly, on a problem of your choosing.

Six phases, and the order that keeps you honestFrame the problem and audit the dataSplit first, then build the pipelineBaseline before any real modelEvaluate: threshold by cost, then segmentPackage, serve, and write it down
The reader has four minutes — every phase above exists so the README can state a number they can trust.

What you will produce

ArtefactWhat good looks like
Problem statementTwo paragraphs: the decision, who makes it, what an error costs
Data auditRow counts, missingness, target balance, leakage checks performed
Training codeRuns start to finish from a clean checkout with one command
Saved pipelineOne file containing preprocessing and model together
EvaluationBaseline, chosen metric, held-out score, per-segment breakdown
Error analysisTwenty worst predictions inspected, with a written hypothesis
Serving layerAn API or app that loads the pipeline and returns predictions
Model cardIntended use, training data, metrics, limitations, ethical notes
READMEProblem, result, how to run it, what you would do next

Choosing a problem you will not regret

More capstones fail at the choice than at the modelling. Apply four tests before committing.

Can you state the decision? Complete this sentence: "When the model outputs X, someone does Y instead of Z." If you cannot, you have a dataset, not a problem.

Is the target real and available before the decision? Every feature must be knowable at prediction time. This kills more datasets than people expect.

Is there enough data? Roughly 1,000 rows minimum for a simple problem, and for classification at least 100 examples of the rarest class you care about. Fewer than that and your evaluation is noise.

Is it boring enough to finish? A tabular regression or binary classification you complete beats an ambitious multi-modal system you abandon at 60%.

ProblemTypeWhy it worksWatch out for
House price predictionRegressionRich features, intuitive, skewed target teaches log transformsOutlier mansions dominating RMSE
Telecom churnClassificationClear business action, moderate imbalanceFeatures recorded after cancellation
Loan defaultClassificationAsymmetric costs make threshold tuning meaningfulFairness across demographic groups
Bike-share demandRegressionStrong seasonal features, easy to engineerMust split by time, never randomly
Hospital readmissionClassificationHeavy imbalance forces proper metricsLeakage from discharge-time fields
Employee attritionClassificationSmall, clean, fast to iterateOften too small for a reliable test set

Avoid datasets where the winning approach is public and copy-pasteable. You will learn more from a messy dataset with an 0.72 AUC than from one where everybody scores 0.99.

Phase one: frame it, then audit the data

Write the problem statement before loading anything. Half a page:

  • The decision being supported, and by whom.
  • The prediction point — the exact moment the model is called.
  • What a false positive costs. What a false negative costs. Numbers if you can get them, a ranking if you cannot.
  • The metric that follows from those costs, and why not accuracy.
  • What "good enough to be useful" would be.

That last item is the one people skip and later regret. Deciding in advance that anything above 0.75 recall at 20% precision is worth shipping stops you from moving the goalposts once you see your results.

Then audit. Not exploratory plotting for its own sake — specific questions with recorded answers.

Python
import pandas as pddf = pd.read_csv("data/raw/dataset.csv")print("shape:", df.shape)print("duplicates:", df.duplicated().sum())print("\nmissing (top 10):")print((df.isna().mean() * 100).sort_values(ascending=False).head(10).round(1))print("\ntarget balance:")print(df[TARGET].value_counts(normalize=True).round(4))print("\nsuspicious values:")print(df.describe().T[["min", "max"]])# the leakage screen: any single feature that predicts the target too wellfrom sklearn.feature_selection import mutual_info_classifmi = mutual_info_classif(df[NUMERIC].fillna(-1), df[TARGET])print(pd.Series(mi, index=NUMERIC).sort_values(ascending=False).head(10))

Record what you find in a short document. Two findings deserve explicit attention:

Sentinel values. Ages of 0 or 999, incomes of −1, dates of 1900-01-01. These encode "unknown" and must become NaN before any imputation, or you will average a 999 into your fill value.

Anything that predicts too well. A single feature with an AUC above 0.95 on its own is leakage until proven otherwise. Investigate before celebrating — this is the failure that makes a capstone worthless while looking spectacular.

Phase two: split first, then build the pipeline

Split before you clean anything, and choose the split type deliberately.

SplitUse whenWhat goes wrong otherwise
Stratified randomIndependent rows, imbalanced targetTest set gets a different class rate; scores become noise
GroupedMultiple rows per customer, patient, deviceThe same entity appears in train and test; the model memorises it
Time-basedAnything predicting forward in timeModel trains on the future to predict the past
Python
from sklearn.model_selection import train_test_splitX, y = df.drop(columns=[TARGET]), df[TARGET]X_temp, X_test, y_temp, y_test = train_test_split(    X, y, test_size=0.2, stratify=y, random_state=42)X_test.assign(**{TARGET: y_test}).to_parquet("data/holdout.parquet")# and do not open it again until the very end

Then put every transformation inside a pipeline so that nothing is ever fitted on data it should not see.

Python
from sklearn.pipeline import Pipelinefrom sklearn.compose import ColumnTransformerfrom sklearn.impute import SimpleImputerfrom sklearn.preprocessing import StandardScaler, OneHotEncoderpre = ColumnTransformer([    ("num", Pipeline([("impute", SimpleImputer(strategy="median")),                      ("scale", StandardScaler())]), NUMERIC),    ("cat", Pipeline([("impute", SimpleImputer(strategy="most_frequent")),                      ("encode", OneHotEncoder(handle_unknown="ignore",                                               min_frequency=10))]), CATEGORICAL),])

min_frequency=10 groups rare categories together, which stops a category seen twice in training from producing a column the model cannot learn anything about. handle_unknown="ignore" stops an unseen category crashing the service later.

Feature engineering is where the marks are

Algorithm choice is worth a few percent. Features are worth more. Spend real time here and document each one.

Python
def add_features(df: pd.DataFrame) -> pd.DataFrame:    out = df.copy()    out["debt_to_income"] = out["debt"] / out["income"].clip(lower=1)    out["spend_per_product"] = out["annual_spend"] / out["num_products"]    out["days_since_contact"] = (        pd.Timestamp("2026-01-01") - pd.to_datetime(out["last_contact"])    ).dt.days    out["is_new"] = (out["tenure_years"] < 1).astype(int)    return out

Two rules. Ratios and differences almost always beat their raw components, because they encode a relationship the model would otherwise have to discover. And any feature built from statistics across rows — a group mean, a target encoding — must be computed inside cross-validation folds, or it leaks.

Phase three: baseline first, then models

Compute the dumb baseline before anything else, so that every later number has a reference point.

Python
from sklearn.dummy import DummyClassifierfrom sklearn.model_selection import cross_val_score, StratifiedKFoldcv = StratifiedKFold(5, shuffle=True, random_state=42)dummy = DummyClassifier(strategy="prior")print("baseline:", cross_val_score(dummy, X_temp, y_temp, cv=cv,                                   scoring="average_precision").mean())

Then work up in complexity, recording every result in a table you keep as you go.

Python
from sklearn.linear_model import LogisticRegressionfrom sklearn.ensemble import RandomForestClassifier, HistGradientBoostingClassifiercandidates = {    "logistic": LogisticRegression(max_iter=2000, class_weight="balanced"),    "random_forest": RandomForestClassifier(n_estimators=400, n_jobs=-1,                                            class_weight="balanced_subsample",                                            random_state=42),    "hist_gbm": HistGradientBoostingClassifier(max_iter=400, random_state=42),}rows = []for name, clf in candidates.items():    pipe = Pipeline([("pre", pre), ("clf", clf)])    s = cross_val_score(pipe, X_temp, y_temp, cv=cv,                        scoring="average_precision", n_jobs=-1)    rows.append({"model": name, "ap_mean": s.mean(), "ap_std": s.std()})print(pd.DataFrame(rows).sort_values("ap_mean", ascending=False))

Report the standard deviation, not just the mean. A model at 0.61 ± 0.02 is not worse than one at 0.63 ± 0.08: the second one's lead is inside its own noise, and claiming it is the most common way capstones overstate their results.

Tune only the top one or two candidates, with randomised search over a log-spaced grid, and keep the tuning inside cross-validation on the training data.

Phase four: evaluate honestly

Now open the held-out file. Once.

Python
final = best_pipeline.fit(X_temp, y_temp)holdout = pd.read_parquet("data/holdout.parquet")X_hold, y_hold = holdout.drop(columns=[TARGET]), holdout[TARGET]probs = final.predict_proba(X_hold)[:, 1]

Three analyses turn a score into an argument.

Threshold selection against real costs

Choose the threshold without looking at the holdout results, using out-of-fold predictions on the training data — every row predicted by a model that never saw it.

Python
import numpy as npfrom sklearn.model_selection import cross_val_predictoof = cross_val_predict(best_pipeline, X_temp, y_temp, cv=cv,                        method="predict_proba")[:, 1]COST_FN, COST_FP = 220.0, 4.0rows = []for t in np.arange(0.02, 0.99, 0.01):    pred = oof >= t    fn = int(((y_temp == 1) & ~pred).sum())    fp = int(((y_temp == 0) & pred).sum())    rows.append({"threshold": round(t, 2), "fn": fn, "fp": fp,                 "cost": fn * COST_FN + fp * COST_FP})costs = pd.DataFrame(rows)print(costs.nsmallest(5, "cost"))

Set CHOSEN_THRESHOLD from this table, then apply it unchanged to the holdout, and report it as part of the model. A model shipped at 0.5 because that is the default has an unexamined assumption that both errors cost the same.

Segment breakdown

Python
holdout["prob"] = probsholdout["pred"] = (probs >= CHOSEN_THRESHOLD).astype(int)for col in ["region", "age_band", "tenure_band"]:    print(holdout.groupby(col).apply(        lambda g: pd.Series({            "n": len(g),            "positive_rate": g[TARGET].mean(),            "recall": (g[(g[TARGET] == 1)]["pred"]).mean(),            "precision": g[g["pred"] == 1][TARGET].mean(),        }), include_groups=False).round(3))

An aggregate AUC of 0.87 can hide 0.91 on one region and 0.58 on another. If any segment corresponds to a protected characteristic, this table is not optional — it is the difference between a model and a liability.

Error analysis

Sort by absolute error, take the twenty worst, and look at the actual rows.

Python
worst = holdout.assign(err=(holdout[TARGET] - holdout["prob"]).abs()) \               .nlargest(20, "err")print(worst[FEATURES + [TARGET, "prob"]])

Then write a paragraph on what you see. Are they all from one region? All recent sign-ups? All missing the same field? This paragraph is frequently the most impressive thing in a capstone, because it shows you interrogated the model rather than just scoring it — and it usually suggests the next feature.

Phase five: package and serve

Save the whole pipeline, plus metadata, plus a fixture for verification.

Python
import joblib, json, sklearn, sysfrom datetime import datetime, timezonejoblib.dump(final, "artifacts/model_v1.0.0.joblib")json.dump({    "version": "1.0.0",    "trained_at": datetime.now(timezone.utc).isoformat(),    "sklearn": sklearn.__version__,    "python": sys.version.split()[0],    "features": list(X_temp.columns),    "decision_threshold": CHOSEN_THRESHOLD,    "holdout_metrics": {"average_precision": 0.612, "roc_auc": 0.871},    "baseline_average_precision": 0.083,}, open("artifacts/model_v1.0.0.meta.json", "w"), indent=2)fixture = X_hold.head(5)json.dump({"inputs": fixture.to_dict("records"),           "expected": final.predict_proba(fixture)[:, 1].round(6).tolist()},          open("artifacts/fixture.json", "w"), indent=2)

That fixture is checked at service startup. If the loaded model produces different numbers — because someone upgraded a library or deployed the wrong file — the service fails loudly instead of quietly returning wrong answers for a month.

Then the service itself. Keep it small; it is judged on correctness, not features.

Python
from fastapi import FastAPIfrom pydantic import BaseModel, Fieldfrom typing import Literalimport joblib, pandas as pdapp = FastAPI(title="Churn API", version="1.0.0")MODEL = joblib.load("artifacts/model_v1.0.0.joblib")THRESHOLD = 0.34class Features(BaseModel):    income: float = Field(ge=0, le=1_000_000)    age: int = Field(ge=18, le=100)    tenure_years: float = Field(ge=0, le=60)    num_products: int = Field(ge=1, le=10)    region: Literal["north", "south", "east", "west"]@app.get("/health")def health() -> dict:    return {"status": "ok", "model_version": "1.0.0"}@app.post("/predict")def predict(f: Features) -> dict:    p = float(MODEL.predict_proba(pd.DataFrame([f.model_dump()]))[0, 1])    return {"probability": round(p, 4), "will_churn": p >= THRESHOLD,            "threshold": THRESHOLD, "model_version": "1.0.0"}

Add a Dockerfile with pinned dependencies and confirm the container starts and answers a request. The README should contain the exact docker build, docker run, and curl commands, and they should work when pasted.

Phase six: write it down

The model card is a page and takes forty minutes.

SectionContents
Intended useThe decision it supports; who uses it; how often
Out of scopeUses it must not be put to — credit decisions, hiring, anything legal
Training dataSource, date range, row count, known biases in collection
FeaturesList, with anything ethically sensitive flagged
PerformanceMetric, held-out value, baseline, per-segment table
LimitationsWhere it fails; the extrapolation range; what would degrade it
MonitoringWhat to track and what should trigger retraining

Write the limitations section as if a sceptical colleague will read it, because stating a weakness plainly reads as competence while omitting it reads as not having looked.

The mistakes that sink these projects

MistakeHow it shows upPrevention
LeakageAUC of 0.99 that nobody questionsAsk of every feature: known before the decision?
Test set used repeatedlyReported score is the best of many attemptsWrite the holdout to a file and leave it shut
Preprocessing outside the pipelineServing code re-implements scaling and driftsEverything in Pipeline, saved as one object
Accuracy on imbalanced data96% accuracy, zero positives caughtAverage precision or balanced accuracy
No baselineNo way to tell whether the model helpsDummyClassifier in the first hour
Tuning before featuresDays of grid search for 0.4%Feature work first; tune once, at the end
Random split on time-series dataExcellent scores, useless modelSplit by date
Only aggregate metricsFails badly on a segment nobody checkedPer-segment table
Nothing runs from a clean checkoutReviewer cannot reproduce anythingTest it in a fresh clone before submitting

A schedule that finishes

StageShare of effortDone when
Framing and data audit15%Problem statement written; leakage screen run
Split and pipeline10%Holdout saved; pipeline transforms without error
Features and models30%Comparison table with means and standard deviations
Evaluation and error analysis20%Threshold chosen; segments checked; 20 errors read
Packaging and serving15%curl against the container returns a prediction
Writing up10%README and model card complete

The temptation is to spend 60% on modelling. Resist it. The modelling is the part most easily done well with least effort, and the parts people skip — framing, error analysis, packaging, writing — are exactly the parts that distinguish a finished system from a notebook.

What this means when you build something

Work so that every claim you make is checkable. That means the holdout file stays closed until the end, the baseline is computed first so the headline number has a reference, the pipeline is one saved object so nobody can ask whether the serving code matches training, and the fixture test proves the deployed artefact behaves like the trained one.

Then be honest in the write-up. A project reporting 0.61 average precision against a 0.08 baseline, with a clear account of where it fails and which segment it underserves, is stronger than one reporting 0.99 with no baseline and no error analysis — because the first is a system somebody could deploy, and the second is a number nobody can trust.