Course Content
MLOps for AI
4 sections · 9 lessons
Capstone Project: End-to-End MLOps Pipeline for Churn Prediction
Here is the brief you are working to. A telecoms company has 10,000 customers on file. Roughly 26.5% of them cancel within a year. Marketing wants a list, refreshed weekly, of who is about to leave, so they can send a retention offer that costs USD 20 per customer and succeeds about 30% of the time. A retained customer is worth USD 400 in remaining margin.
Most people who attempt this produce a notebook. It loads the CSV, engineers some features, trains a gradient-boosted tree, prints AUC: 0.84, and stops. That notebook is worth almost nothing to the company, because marketing cannot run it, nobody can tell which version produced last week's list, and when it starts performing badly no one will notice.
The gap between that notebook and a system is what you are going to close. Four parts, each producing an artefact the next one consumes:
| Part | You build | It produces | Consumed by |
|---|---|---|---|
| 1. Track and register | Multi-run training with MLflow, a promotion gate | A registered model version with a @champion alias | The service, which resolves the alias |
| 2. Package | A Flask inference service and a hardened Dockerfile | An image identified by digest | The deployment manifest |
| 3. Deploy | Deployment, Service, HPA manifests | A running, scaling, health-checked endpoint | Real traffic |
| 4. Observe | Prometheus metrics, alerts, documentation | Evidence the thing still works | The next retraining decision |
The test of this project is not the AUC. It is whether a colleague who has never seen the repository can find out what is deployed, why it was chosen, and how to undo it — using only what you built.
Part 1 — Generate data you can reason about
Use synthetic data with a known structure. Real datasets hide their generating process; synthetic data lets you verify that your pipeline recovers signal you deliberately planted, which is how you catch a broken feature pipeline early.
1import numpy as np, pandas as pd23def make_customers(n: int = 10_000, seed: int = 42) -> pd.DataFrame:4 rng = np.random.default_rng(seed)5 tenure = rng.integers(1, 73, n)6 monthly = np.round(rng.normal(70, 25, n).clip(18.5, 120), 2)7 contract = rng.choice(["month-to-month", "one-year", "two-year"],8 n, p=[0.55, 0.24, 0.21])9 support_calls = rng.poisson(1.2, n)10 has_fibre = rng.random(n) < 0.441112 # Planted signal: a log-odds model we can check the trained model recovers.13 logit = (-1.9714 - 0.030 * tenure15 + 0.011 * monthly16 + 0.34 * support_calls17 + 0.42 * has_fibre18 + np.where(contract == "month-to-month", 1.15,19 np.where(contract == "one-year", 0.10, -0.85)))20 churn = rng.random(n) < 1 / (1 + np.exp(-logit))2122 return pd.DataFrame({23 "customer_id": [f"C{i:06d}" for i in range(n)],24 "tenure_months": tenure, "monthly_charges": monthly,25 "contract": contract, "support_calls": support_calls,26 "has_fibre": has_fibre, "churned": churn.astype(int),27 })2829df = make_customers()30print(len(df), df["churned"].mean().round(4)) # 10000 0.266Two coefficients are worth remembering for later: contract type dominates (a swing of 2.0 in log-odds between two-year and month-to-month), and tenure protects (−0.030 per month, so 24 months of tenure is −0.72). If your trained model's feature importances do not put contract and tenure_months at the top, something in your pipeline is broken — a shuffled join, a leaked column, a mis-encoded category. Synthetic data gives you that check for free.
Train several candidates, log all of them
1import mlflow, mlflow.sklearn, subprocess, hashlib2from mlflow.models import infer_signature3from sklearn.compose import ColumnTransformer4from sklearn.ensemble import RandomForestClassifier, GradientBoostingClassifier5from sklearn.linear_model import LogisticRegression6from sklearn.metrics import roc_auc_score, average_precision_score7from sklearn.model_selection import train_test_split8from sklearn.pipeline import Pipeline9from sklearn.preprocessing import OneHotEncoder, StandardScaler1011NUM = ["tenure_months", "monthly_charges", "support_calls"]12CAT = ["contract", "has_fibre"]1314def build(estimator) -> Pipeline:15 pre = ColumnTransformer([16 ("num", StandardScaler(), NUM),17 ("cat", OneHotEncoder(handle_unknown="infrequent_if_exist",18 min_frequency=20), CAT),19 ])20 return Pipeline([("pre", pre), ("clf", estimator)])2122X = df.drop(columns=["churned", "customer_id"])23y = df["churned"]24X_tr, X_te, y_tr, y_te = train_test_split(25 X, y, test_size=0.2, stratify=y, random_state=42)2627data_hash = hashlib.sha256(28 pd.util.hash_pandas_object(df, index=True).values.tobytes()).hexdigest()[:16]29git_sha = subprocess.check_output(["git", "rev-parse", "HEAD"], text=True).strip()3031CANDIDATES = {32 "logreg": LogisticRegression(max_iter=1000, C=1.0, random_state=42),33 "rf": RandomForestClassifier(n_estimators=400, max_depth=8,34 min_samples_leaf=20, random_state=42),35 "gbm": GradientBoostingClassifier(n_estimators=300, max_depth=3,36 learning_rate=0.05, random_state=42),37}3839mlflow.set_experiment("churn-prediction")4041for name, est in CANDIDATES.items():42 with mlflow.start_run(run_name=name):43 pipe = build(est).fit(X_tr, y_tr)44 prob = pipe.predict_proba(X_te)[:, 1]4546 mlflow.log_params({"model_type": name, **{47 f"clf.{k}": v for k, v in est.get_params().items()48 if isinstance(v, (int, float, str, bool)) or v is None}})49 mlflow.set_tags({"git_sha": git_sha, "data_hash": data_hash,50 "rows": len(df)})51 mlflow.log_metrics({52 "test_auc": roc_auc_score(y_te, prob),53 "test_ap": average_precision_score(y_te, prob),54 "test_pos_rate": float(y_te.mean()),55 })56 # Slice metrics: the average hides the segments marketing cares about.57 for label, mask in {58 "month_to_month": X_te["contract"] == "month-to-month",59 "new_customers": X_te["tenure_months"] <= 6,60 "high_value": X_te["monthly_charges"] > 90,61 }.items():62 m = mask.to_numpy()63 if m.sum() >= 100 and y_te[m].nunique() == 2:64 mlflow.log_metric(f"slice.{label}.auc",65 roc_auc_score(y_te[m], prob[m]))66 mlflow.log_metric(f"slice.{label}.n", int(m.sum()))6768 mlflow.sklearn.log_model(69 sk_model=pipe, name="model",70 signature=infer_signature(X_tr, pipe.predict_proba(X_tr)),71 input_example=X_tr.iloc[:3],72 pyfunc_predict_fn="predict_proba") # the service needs probabilitiesLogging the whole Pipeline rather than the bare estimator is the decision that prevents the most common bug in this project. If you log only the classifier, your serving code has to reimplement the scaling and one-hot encoding — and the day someone adds a category or changes the scaler, training and serving drift apart silently. One object in, one object out.
The pyfunc_predict_fn="predict_proba" argument matters in Part 2. The service loads the model through MLflow's generic pyfunc interface, and without that argument its predict returns hard 0/1 labels rather than the probabilities the threshold is applied to.
Choose a threshold with money, not with F1
AUC ranks candidates. It does not tell marketing whom to call. That needs a threshold, and the right threshold comes from the economics in the brief.
The test set has 2,000 customers, of whom 532 churn. Score them with the GBM (test AUC about 0.77; the planted signal is deliberately noisy) and count what happens at four thresholds. These are the numbers the code above produces with scikit-learn 1.9; other versions may differ slightly.
| Threshold | Flagged | True churners caught | Offer cost | Expected retained (30%) | Margin saved | Net |
|---|---|---|---|---|---|---|
| 0.50 | 263 | 167 | USD 5,260 | 50.1 | USD 20,040 | USD 14,780 |
| 0.30 | 700 | 335 | USD 14,000 | 100.5 | USD 40,200 | USD 26,200 |
| 0.15 | 1,310 | 480 | USD 26,200 | 144 | USD 57,600 | USD 31,400 |
| 0.05 | 1,805 | 524 | USD 36,100 | 157.2 | USD 62,880 | USD 26,780 |
Check the arithmetic on the 0.15 row: 1,310 offers × USD 20 = USD 26,200 spent; 480 true churners × 0.30 retention success = 144 customers kept; 144 × USD 400 = USD 57,600; net USD 31,400.
The default threshold of 0.5 leaves 31,400 − 14,780 = USD 16,620 per cycle on the table — because it optimises for being right, and the business is paying for churners caught. Note also that 0.05 is worse than 0.15 despite catching more churners: the extra 495 offers cost USD 9,900 and reach only 44 more churners, worth 44 × 0.30 × 400 = USD 5,280. There is a maximum, and you can predict roughly where it sits before running anything. An offer pays for itself when P(churn) × 0.30 × 400 is more than 20, that is when P(churn) is above 20 / 120 ≈ 0.17. For a reasonably calibrated model the best threshold therefore sits near 0.17, far from any default; the sweep below finds a flat optimum between 0.14 and 0.18, and this project uses 0.15, a round number inside it.
The model is chosen by its ranking quality; the threshold is chosen by the cost of being wrong. Only the first of those two ever appears on a leaderboard.
1import numpy as np23def best_threshold(y_true, prob, offer_cost=20, margin=400, success=0.30):4 rows = []5 for t in np.arange(0.05, 0.95, 0.01):6 flagged = prob >= t7 tp = int((flagged & (y_true == 1)).sum())8 net = tp * success * margin - int(flagged.sum()) * offer_cost9 rows.append((round(float(t), 2), int(flagged.sum()), tp, round(net, 2)))10 return max(rows, key=lambda r: r[3])Log the chosen threshold as a parameter of the run. A probability without its decision threshold is not a decision, and a metric computed at an unrecorded threshold is not evidence.
Register the champion
1from mlflow import MlflowClient23client = MlflowClient()4runs = mlflow.search_runs(experiment_names=["churn-prediction"],5 filter_string=f"tags.data_hash = '{data_hash}'",6 order_by=["metrics.test_auc DESC"])7best = runs.iloc[0]89mv = mlflow.register_model(f"runs:/{best.run_id}/model", "churn-predictor")10for k, v in {"git_sha": git_sha, "data_hash": data_hash,11 "test_auc": f"{best['metrics.test_auc']:.4f}",12 "threshold": "0.15",13 "worst_slice_auc": f"{min(best[c] for c in runs.columns if c.startswith('metrics.slice.') and c.endswith('.auc')):.4f}"}.items():14 client.set_model_version_tag("churn-predictor", mv.version, k, v)1516# Remember what we are replacing BEFORE we replace it.17try:18 prev = client.get_model_version_by_alias("churn-predictor", "champion")19 client.set_registered_model_alias("churn-predictor", "previous-champion",20 prev.version)21except mlflow.exceptions.MlflowException:22 pass23client.set_registered_model_alias("churn-predictor", "champion", mv.version)The data_hash filter on the search is not optional. Ranking runs across different data snapshots produces a leaderboard whose winner is whoever got the easiest split, and it looks exactly like a result.
The previous-champion alias set one line before promotion is what turns a future rollback from an investigation into a lookup.
Part 2 — A service, and an image worth deploying
1# src/churn/api.py2import os, threading, time, logging3import mlflow, pandas as pd4from flask import Flask, request, jsonify5from prometheus_client import Counter, Histogram, Gauge, generate_latest67MODEL_NAME = os.environ.get("MODEL_NAME", "churn-predictor")8ALIAS = os.environ.get("MODEL_ALIAS", "champion")9THRESHOLD = float(os.environ.get("THRESHOLD", "0.15"))1011REQUESTS = Counter("churn_requests_total", "Requests", ["version", "status"])12LATENCY = Histogram("churn_latency_seconds", "Latency", ["version"],13 buckets=(.005, .01, .025, .05, .1, .25, .5, 1, 2.5))14SCORE = Histogram("churn_score", "Predicted probability",15 buckets=[i / 20 for i in range(21)])16READY = Gauge("churn_ready", "1 when a model is loaded")1718app = Flask(__name__)19log = logging.getLogger("churn")2021REQUIRED = ["tenure_months", "monthly_charges", "contract",22 "support_calls", "has_fibre"]232425class ModelHolder:26 def __init__(self):27 self._lock = threading.RLock()28 self.model = None29 self.version = "none"30 self.refresh()31 threading.Thread(target=self._poll, daemon=True).start()3233 def refresh(self):34 c = mlflow.MlflowClient()35 mv = c.get_model_version_by_alias(MODEL_NAME, ALIAS)36 if mv.version == self.version:37 return38 loaded = mlflow.pyfunc.load_model(f"models:/{MODEL_NAME}@{ALIAS}")39 with self._lock: # swap only after a clean load40 self.model, self.version = loaded, mv.version41 READY.set(1)42 log.info("loaded %s version %s", MODEL_NAME, mv.version)4344 def _poll(self):45 while True:46 time.sleep(60)47 try:48 self.refresh()49 except Exception as exc:50 log.warning("refresh failed, keeping v%s: %s", self.version, exc)5152 def predict(self, frame):53 with self._lock:54 return self.model.predict(frame), self.version555657holder = ModelHolder()585960@app.get("/health")61def health():62 return jsonify(status="ok"), 200 # process is alive636465@app.get("/ready")66def ready():67 ok = holder.model is not None68 return jsonify(ready=ok, version=holder.version), (200 if ok else 503)697071@app.post("/predict")72def predict():73 payload = request.get_json(silent=True)74 if payload is None:75 REQUESTS.labels(holder.version, "bad_request").inc()76 return jsonify(error="body must be JSON"), 4007778 rows = payload if isinstance(payload, list) else [payload]79 if len(rows) > 1000:80 REQUESTS.labels(holder.version, "too_large").inc()81 return jsonify(error="max 1000 rows per request"), 4138283 missing = [f for f in REQUIRED if f not in rows[0]]84 if missing:85 REQUESTS.labels(holder.version, "bad_request").inc()86 return jsonify(error=f"missing fields: {missing}"), 4008788 try:89 with LATENCY.labels(holder.version).time():90 proba, version = holder.predict(pd.DataFrame(rows)[REQUIRED])91 scores = [float(p[1]) for p in proba]92 except Exception:93 REQUESTS.labels(holder.version, "error").inc()94 log.exception("prediction failed")95 return jsonify(error="internal error"), 5009697 for s in scores:98 SCORE.observe(s)99 REQUESTS.labels(version, "ok").inc()100 return jsonify(model_version=version, threshold=THRESHOLD,101 predictions=[{"score": round(s, 6),102 "flag": s >= THRESHOLD} for s in scores])103104105@app.get("/metrics")106def metrics():107 return generate_latest(), 200, {"Content-Type": "text/plain"}Four things in that file exist because of specific production failures.
/healthand/readyare different endpoints. Health means the process is alive; readiness means it can serve. A pod reloading its model should leave the load-balancer rotation without being restarted, and conflating the two makes that impossible.- The 1,000-row cap. This is the sale-day OOM in miniature. An unbounded batch endpoint lets one caller allocate arbitrary memory and take the pod down for everyone.
- Every response carries
model_version. Without it you cannot attribute a change in downstream behaviour to a deployment. - The generic 500 body. Returning a stack trace to a caller leaks file paths, library versions and sometimes data. Log the detail; return nothing.
# docker/serve.DockerfileFROM python:3.11-slim@sha256:2f4c9d3b71e08a5c6d4bf2190ac83e7d5b6f01c2a9e4 AS builderWORKDIR /buildRUN apt-get update && apt-get install -y --no-install-recommends build-essential \ && rm -rf /var/lib/apt/lists/*COPY requirements-serve.txt .RUN pip install --no-cache-dir --prefix=/install -r requirements-serve.txtFROM python:3.11-slim@sha256:2f4c9d3b71e08a5c6d4bf2190ac83e7d5b6f01c2a9e4WORKDIR /appCOPY --from=builder /install /usr/localCOPY src/churn/ ./churn/RUN useradd --create-home --uid 1001 appuserUSER appuserENV PYTHONUNBUFFERED=1 OMP_NUM_THREADS=2 THRESHOLD=0.15EXPOSE 8000CMD ["gunicorn", "-w", "2", "-b", "0.0.0.0:8000", "--timeout", "60", \ "--access-logfile", "-", "churn.api:app"]Copy requirements-serve.txt before the source so the dependency layer stays cached when only code changes — that is the difference between an 8-second rebuild and a 6-minute one. OMP_NUM_THREADS=2 matches the CPU limit set in the deployment; leave it unset and the BLAS library spawns one thread per host core, burns the container's CPU quota in a fraction of each scheduling period, and produces latency plateaus that no profiler explains.
Never run flask run in production. It is a single-threaded development server with no timeout handling. Gunicorn with two workers gives you process isolation, so one wedged request does not stall the other worker.
Part 3 — Deploy it so it survives Monday
apiVersion: apps/v1kind: Deploymentmetadata: {name: churn-api, namespace: ml-production}spec: replicas: 3 strategy: rollingUpdate: {maxSurge: 1, maxUnavailable: 0} selector: {matchLabels: {app: churn-api}} template: metadata: labels: {app: churn-api} annotations: {prometheus.io/scrape: "true", prometheus.io/port: "8000"} spec: containers: - name: api image: ghcr.io/acme/churn-api@sha256:8b2e1f... ports: [{containerPort: 8000}] env: - {name: MODEL_ALIAS, value: champion} - {name: THRESHOLD, value: "0.15"} - name: MLFLOW_TRACKING_URI valueFrom: {configMapKeyRef: {name: ml-config, key: tracking_uri}} resources: requests: {cpu: "500m", memory: "1Gi"} limits: {cpu: "1000m", memory: "2Gi"} startupProbe: httpGet: {path: /ready, port: 8000} failureThreshold: 24 periodSeconds: 5 readinessProbe: httpGet: {path: /ready, port: 8000} periodSeconds: 10 livenessProbe: httpGet: {path: /health, port: 8000} periodSeconds: 20 lifecycle: {preStop: {exec: {command: ["sleep", "10"]}}} terminationGracePeriodSeconds: 45---apiVersion: v1kind: Servicemetadata: {name: churn-api, namespace: ml-production}spec: selector: {app: churn-api} ports: [{port: 80, targetPort: 8000}]---apiVersion: autoscaling/v2kind: HorizontalPodAutoscalermetadata: {name: churn-api, namespace: ml-production}spec: scaleTargetRef: {apiVersion: apps/v1, kind: Deployment, name: churn-api} minReplicas: 3 maxReplicas: 15 metrics: - type: Resource resource: {name: cpu, target: {type: Utilization, averageUtilization: 65}} behavior: scaleUp: {stabilizationWindowSeconds: 0, policies: [{type: Percent, value: 100, periodSeconds: 30}]} scaleDown: {stabilizationWindowSeconds: 300, policies: [{type: Percent, value: 25, periodSeconds: 60}]}The startupProbe grants 24 × 5 = 120 seconds to pull the model from the registry before liveness checking begins. Omit it and a slow registry means the liveness probe kills the pod within about a minute, every time, forever — a crash loop with no error in the logs beyond "connection refused".
maxUnavailable: 0 keeps all three replicas serving throughout a rollout. preStop: sleep 10 gives endpoint removal time to propagate before the process exits, so a deploy does not produce a scatter of connection-refused errors.
Verify with the checks that would actually catch a broken deploy, not just kubectl get pods:
1kubectl -n ml-production rollout status deployment/churn-api --timeout=300s23# The version the pods actually loaded should equal the registry's champion.4kubectl -n ml-production port-forward svc/churn-api 8080:80 &5curl -s localhost:8080/ready | python -m json.tool67# A real prediction, and its arithmetic sanity: a two-year contract customer8# with long tenure should score well below a new month-to-month one.9curl -s -X POST localhost:8080/predict -H 'Content-Type: application/json' \10 -d '[{"tenure_months":60,"monthly_charges":45.0,"contract":"two-year",11 "support_calls":0,"has_fibre":false},12 {"tenure_months":2,"monthly_charges":99.0,"contract":"month-to-month",13 "support_calls":4,"has_fibre":true}]' | python -m json.toolA rollback you have never performed is a rollback you do not have. Rehearse it in staging before you need it at three in the morning.
Part 4 — Evidence that it still works
The metrics are already exported by the service. What remains is deciding what wakes someone up.
groups:- name: churn-api rules: - alert: ChurnApiErrorRate expr: | sum(rate(churn_requests_total{status="error"}[5m])) / sum(rate(churn_requests_total[5m])) > 0.02 for: 5m labels: {severity: page} - alert: ChurnApiLatencyP99 expr: | histogram_quantile(0.99, sum by (le) (rate(churn_latency_seconds_bucket[5m]))) > 0.5 for: 10m labels: {severity: page} - alert: ChurnScoreDrift expr: | abs(rate(churn_score_sum[1h]) / rate(churn_score_count[1h]) - 0.265) > 0.06 for: 30m labels: {severity: ticket} - alert: ChurnFlagRateAnomaly expr: | sum(rate(churn_score_bucket{le="0.15"}[1h])) / sum(rate(churn_score_count[1h])) < 0.20 for: 1h labels: {severity: ticket}ChurnScoreDrift is the alert that distinguishes this from an ordinary web service. The training base rate is 0.265, so a well-behaved model scoring representative traffic should average close to that. If the mean predicted probability moves to 0.35, something upstream has changed — a feature's units, a new contract type encoded as unknown, a data pipeline that started sending nulls. Nothing errors. Latency is fine. Only the score distribution shows it.
Note the severities: operational failures page, distributional shifts open a ticket. Paging on drift at 03:00 for something that cannot be fixed until morning is how teams learn to ignore pages.
Documentation that is actually used
Write four short documents, not one long one, because they have four different readers.
| Document | Reader | Must answer |
|---|---|---|
README.md | A new engineer | How do I run this locally in under ten minutes? |
MODEL_CARD.md | Marketing, compliance | What does it predict, on what data, how accurate per segment, what are its limits? |
API.md | The team calling the endpoint | Request shape, response shape, every error code, rate limits |
RUNBOOK.md | Whoever is on call at 03:00 | For each alert: what it means, first three things to check, exact rollback command |
The runbook is the one most often skipped and the only one read under stress. Its rollback entry should be a command that can be copied, not a description of a procedure:
1# Roll back the model (no rebuild, no redeploy - pods reload within 60s)2python -c "3from mlflow import MlflowClient4c = MlflowClient()5prev = c.get_model_version_by_alias('churn-predictor', 'previous-champion')6bad = c.get_model_version_by_alias('churn-predictor', 'champion')7c.set_registered_model_alias('churn-predictor', 'champion', prev.version)8c.set_model_version_tag('churn-predictor', bad.version, 'blocked', 'true')9print(f'champion: {bad.version} -> {prev.version}')10"1112# Roll back the SERVICE (code or config problem, not a model problem)13kubectl -n ml-production rollout undo deployment/churn-apiTwo different rollbacks for two different failures. Confusing them wastes the first ten minutes of every incident.
Judging your own work
Grade yourself against what the system can do, not what you wrote. Each row below is either demonstrably true or it is not.
| Claim | How to prove it |
|---|---|
| Every deployed model traces to a run | From the live model_version, recover git SHA, data hash and metrics without opening a notebook |
| The threshold is justified | Point to the expected-value table and the parameter recording it |
| Rollback works | Actually do it in a staging namespace and time it — under two minutes |
| Serving matches training | A test asserting the API's score equals the training pipeline's for the same ten rows, to 1e-6 |
| A bad input fails cleanly | POST a missing field, a null, a 5,000-row batch, an unknown contract type — 400 or 413, never 500 |
| Degradation is visible | Send traffic with all contracts set to month-to-month; the drift alert should fire |
| A colleague can run it | Someone else follows the README on a clean machine without asking you anything |
The mistakes that show up in almost every attempt
| Mistake | Why it happens | What it costs |
|---|---|---|
| Logging the estimator, not the whole pipeline | The classifier feels like "the model" | Serving reimplements preprocessing; the two drift apart silently |
| Hardcoding a model version in the manifest | It is the obvious thing to type | Rollback needs a code change, review and deploy instead of one metadata write |
| Using threshold 0.5 | It is the default | USD 16,620 per cycle in this project's economics |
| One probe endpoint for liveness and readiness | They look the same | Pods restarted mid-reload; crash loops during model refresh |
No startupProbe | Nothing fails until the model gets big | Permanent crash loop with no useful error |
| Unbounded batch endpoint | Nobody sends 5,000 rows in testing | One caller OOMs the pod for everyone |
| Comparing runs across data versions | The leaderboard sorts by AUC and looks authoritative | Promoting whichever run got the easiest split |
No previous-champion pointer | You only need it once | Rollback becomes an archaeology exercise during an outage |
What to build first if you only have a day
If you cannot complete all four parts, the order that produces the most working system per hour is not the order they are listed in. Build the service and its rollback path first, against a model you register in ten minutes with default hyperparameters. A mediocre model you can deploy, observe and revert is a system. An excellent model in a notebook is a screenshot.
Then add the score histogram and the drift alert, because they are what tell you whether anything you do next matters. Then improve the model — and now every improvement is measurable, promotable and reversible, which means you can move quickly precisely because you can undo things.
The instinct to perfect the model first is strong and it is backwards. In the brief you were given, the difference between AUC 0.77 and 0.79 is worth a few hundred dollars a cycle. The difference between a two-minute rollback and a four-hour one, on the day something goes wrong, is worth an order of magnitude more.