Course Content
MLOps for AI
4 sections · 9 lessons
MLflow, DVC, and Kubeflow Basics
A team of four has been working on a credit-risk model for six weeks. Their record of it is a spreadsheet, experiments_v4.xlsx, with 63 rows and columns headed lr, depth, auc, notes. Row 41 reads 0.912 — best so far!!. Nobody can reproduce it.
The reasons are worth listing because each one points at a different tool:
- The row does not say which commit produced it, or what the other eleven hyperparameters were set to. Somebody typed four columns and left the rest.
- It says nothing about the data. Two team members were using
data/train_v7.csv; one had regenerated it after a bug fix and kept the same filename. Their AUCs are not comparable and nobody realised. - The 0.912 run took nine hours on one person's workstation. It cannot be rerun anywhere else, because "run it" means opening a laptop and executing cells in the right order.
Three failures, three different shapes. One is about recording what happened. One is about versioning large files that cannot go in Git. One is about running work reproducibly on shared compute. MLflow, DVC and Kubeflow each solve one of them well and the other two badly. Teams that treat them as competitors pick one, discover it does not cover everything, and conclude the tool is bad.
| Question | Tool | What it stores | What it does not do |
|---|---|---|---|
| What did I run, and what came out? | MLflow | Params, metrics, artefacts, model versions in a database | Does not version your data files; does not schedule or provision compute |
| Which exact bytes of data and model did this use? | DVC | Content hashes in Git; bytes in object storage | No metric comparison UI worth the name; no serving |
| Where does this actually execute? | Kubeflow | Containerised pipeline steps on Kubernetes | Not an experiment tracker; heavy for a single script |
These three tools are not alternatives to each other. A team that uses all three uses each for exactly one job, and the seams between them are where the design work lies.
MLflow: a database for things you tried
MLflow's core abstraction is the run: one execution of training code, with everything about it recorded in one place. Runs live inside an experiment, which is just a named folder for related runs.
Each run holds four kinds of thing, and knowing which is which prevents most misuse:
| Category | Type | Immutable? | Examples | Common mistake |
|---|---|---|---|---|
| Params | Key → string | Yes, set once | max_depth=6, data_version=v7 | Trying to log a param twice; logging a metric as a param |
| Metrics | Key → float, with step and timestamp | No, append-only series | val_auc per epoch | Logging only the final value, losing the curve |
| Tags | Key → string | No, mutable | owner=priya, stage=candidate | Not used at all, then search is impossible |
| Artefacts | Files | Yes | Model files, plots, confusion matrices, the config | Logging a 4 GB dataset as an artefact of every run |
The params-are-immutable rule catches people out. If you try to log_param("max_depth", 6) twice with different values in one run, MLflow raises. That is deliberate — a run has one configuration, by definition. If your value changes during training, it is a metric.
Setting up a tracking server
Running with no configuration writes to a local ./mlruns directory, which is fine for the first hour and useless for a team. A shared server needs two backends: a database for the structured records, and object storage for the files.
1pip install "mlflow>=3" psycopg2-binary boto323mlflow server \4 --host 0.0.0.0 --port 5000 \5 --backend-store-uri postgresql://mlflow:secret@db:5432/mlflow \6 --artifacts-destination s3://ml-artifacts/mlflow \7 --serve-artifactsClients then point at it with an environment variable, never a hardcoded string in the training script:
export MLFLOW_TRACKING_URI=http://mlflow.internal:5000Splitting the two stores is the whole reason MLflow scales. Metrics are small and queried constantly, so they belong in Postgres. Model files are large and read rarely, so they belong in object storage. A team that leaves everything on a shared filesystem discovers at about 5,000 runs that listing an experiment takes 40 seconds.
A tracked run, written out
1import mlflow, mlflow.sklearn2from mlflow.models import infer_signature3from sklearn.ensemble import RandomForestClassifier4from sklearn.metrics import roc_auc_score, average_precision_score56mlflow.set_experiment("credit-risk")78params = {"n_estimators": 400, "max_depth": 6, "min_samples_leaf": 20,9 "class_weight": "balanced", "random_state": 42}1011with mlflow.start_run(run_name="rf-depth6-balanced") as run:12 mlflow.log_params(params)13 mlflow.set_tags({"owner": "priya", "data_version": "v7",14 "git_sha": git_sha()})1516 model = RandomForestClassifier(**params).fit(X_train, y_train)1718 proba = model.predict_proba(X_val)[:, 1]19 mlflow.log_metrics({20 "val_auc": roc_auc_score(y_val, proba),21 "val_ap": average_precision_score(y_val, proba),22 "val_pos_rate": float(y_val.mean()),23 })2425 # A signature records the expected input/output schema. Without it,26 # serving accepts a wrongly-ordered dataframe and silently mispredicts.27 signature = infer_signature(X_train, model.predict_proba(X_train))2829 mlflow.sklearn.log_model(30 sk_model=model,31 name="model", # MLflow 2.x called this `artifact_path`32 signature=signature,33 input_example=X_train.iloc[:5],34 )3536 mlflow.log_artifact("configs/train_rf.yaml")37 print("run_id:", run.info.run_id)The signature is the line most people omit and later wish they had not. A scikit-learn pipeline fed a DataFrame with the same columns in a different order will not complain; it will use positional order internally and produce confident nonsense. A logged signature makes the serving layer reject that input instead.
Autologging: the fast path and its limits
1mlflow.sklearn.autolog() # or .xgboost, .pytorch, .lightgbm, .transformers23with mlflow.start_run():4 model = RandomForestClassifier(n_estimators=400, max_depth=6).fit(X_train, y_train)That captures every constructor argument as a param, training metrics, the fitted model, and the signature — with no logging code at all. It is genuinely the right default when you are exploring.
What it does not capture is anything the framework does not know about: which data version you used, what your feature engineering did, which git commit is checked out, and any custom metric such as profit at a chosen threshold. So the practical pattern is autolog plus a handful of explicit calls:
1mlflow.sklearn.autolog(log_datasets=False) # datasets can be enormous23with mlflow.start_run():4 mlflow.set_tags({"git_sha": git_sha(), "data_version": "v7"})5 model = fit(X_train, y_train)6 mlflow.log_metric("profit_at_p30", profit_at_threshold(model, X_val, y_val, 0.30))Wrapping it so the team uses it consistently
Four people logging by hand will produce four naming conventions, and comparison becomes impossible. A thin wrapper enforces the ones that matter.
1import contextlib, json, subprocess2from pathlib import Path3import mlflow45class RunLogger:6 def __init__(self, experiment: str, tracking_uri: str | None = None):7 if tracking_uri:8 mlflow.set_tracking_uri(tracking_uri)9 mlflow.set_experiment(experiment)1011 @contextlib.contextmanager12 def run(self, name: str, config: dict, data_version: str):13 dirty = bool(subprocess.check_output(14 ["git", "status", "--porcelain"], text=True).strip())15 if dirty:16 raise RuntimeError("commit before training; a dirty tree makes git_sha a lie")1718 with mlflow.start_run(run_name=name) as r:19 mlflow.log_params(_flatten(config))20 mlflow.set_tags({21 "git_sha": subprocess.check_output(22 ["git", "rev-parse", "HEAD"], text=True).strip(),23 "data_version": data_version,24 })25 Path("config_snapshot.json").write_text(json.dumps(config, indent=2))26 mlflow.log_artifact("config_snapshot.json")27 yield r2829def _flatten(d: dict, prefix: str = "") -> dict:30 """MLflow params are flat; nested config must be flattened first."""31 out = {}32 for k, v in d.items():33 key = f"{prefix}{k}"34 if isinstance(v, dict):35 out.update(_flatten(v, prefix=f"{key}."))36 else:37 out[key] = v38 return out_flatten is small and essential. MLflow params are a flat key-value map, so a nested config dictionary logged directly becomes one param whose value is the string "{'model': {'max_depth': 6, ...}}". You cannot filter or sort on that. Flattened, it becomes model.max_depth = 6 and the UI can group by it.
Getting things back out
1import mlflow2from mlflow.tracking import MlflowClient34# Query as a DataFrame — the fastest way to compare dozens of runs.5runs = mlflow.search_runs(6 experiment_names=["credit-risk"],7 filter_string="metrics.val_auc > 0.88 and tags.data_version = 'v7'",8 order_by=["metrics.val_auc DESC"],9 max_results=20,10)11print(runs[["run_id", "params.model.max_depth", "metrics.val_auc"]])1213best_id = runs.iloc[0]["run_id"]14model = mlflow.pyfunc.load_model(f"runs:/{best_id}/model")15preds = model.predict(X_test)1617# The client gives you the full metric history, not just the last value.18client = MlflowClient()19for point in client.get_metric_history(best_id, "val_auc"):20 print(point.step, round(point.value, 4))mlflow.pyfunc.load_model is the loader to reach for by default. It returns a uniform predict interface whether the underlying model is scikit-learn, XGBoost or PyTorch, which means downstream code does not have to know or care. Reach for mlflow.sklearn.load_model only when you genuinely need the native object — to inspect feature_importances_, for instance.
DVC: Git for files too big for Git
Git stores every version of every file in the repository. That works beautifully for text and disastrously for a 2.3 GB CSV: the repository grows by 2.3 GB per version, clones take forever, and the delta compression that makes Git efficient does nothing on binary data.
DVC's move is to split the problem. The content goes to object storage. A tiny pointer file containing the content's hash goes into Git. Git then versions the pointers, which is exactly what Git is good at.
1pip install "dvc[s3]"23cd credit-risk4dvc init # creates .dvc/ and stages it in Git5dvc remote add -d storage s3://ml-data/credit-risk6git commit -m "chore: init dvc"78dvc add data/raw/applications.csv # 2.3 GB9git add data/raw/applications.csv.dvc data/raw/.gitignore10git commit -m "data: applications snapshot 2024-04-15"1112dvc push # bytes go to S3The pointer file is about five lines:
outs:- md5: 4b7e2a91c0d83f5e6a1b9c4d7e0f2a38 size: 2451337216 hash: md5 path: applications.csvA filename is a label; a content hash is an identity. Only one of the two survives somebody quietly regenerating the file.
Now a colleague runs git checkout <commit> followed by dvc pull and gets byte-identical data for that commit. The credit-risk team's second failure — two people silently using different train_v7.csv files — becomes impossible, because the filename is no longer the identity. The hash is.
Pipelines: making the dependency graph explicit
DVC's second capability is a stage graph declared in dvc.yaml. Each stage names its command, its dependencies, its parameters and its outputs.
stages: prepare: cmd: python src/prepare.py deps: - src/prepare.py - data/raw/applications.csv params: - prepare.test_size - prepare.seed outs: - data/processed/ train: cmd: python src/train.py deps: - src/train.py - data/processed/ params: - train.n_estimators - train.max_depth outs: - models/model.joblib metrics: - reports/metrics.json: cache: false plots: - reports/roc.csv: x: fpr y: tprWith params.yaml alongside it:
prepare: test_size: 0.2 seed: 42train: n_estimators: 400 max_depth: 6Then dvc repro hashes every dependency and every parameter, compares against the last run, and re-executes only the stages whose inputs changed. Change max_depth from 6 to 8 and prepare is skipped entirely — a real saving when preparation takes forty minutes. Change the raw data and both stages rerun.
1dvc repro # rebuild what is stale2dvc metrics diff HEAD~1 # what did that change do to the numbers?3dvc exp run -S train.max_depth=8 # a throwaway experiment, no commit needed4dvc exp show # table of experiments with params and metricsDVC's Python API lets other code read versioned data without checking it out:
1import dvc.api2import pandas as pd34# Read a specific historical version straight from remote storage.5with dvc.api.open("data/raw/applications.csv",6 repo="https://github.com/acme/credit-risk",7 rev="v1.2.0", mode="r") as f:8 df = pd.read_csv(f)910params = dvc.api.params_show() # the params.yaml for the current commit11url = dvc.api.get_url("models/model.joblib", rev="v1.2.0")Kubeflow: running the work somewhere that is not a laptop
The third failure — a nine-hour run that only exists on one workstation — is a compute problem. Kubeflow is a set of Kubernetes components for machine learning; the relevant piece here is Kubeflow Pipelines, which turns each step of a workflow into a container that Kubernetes schedules.
| Component | Purpose |
|---|---|
| Pipelines (KFP) | Define and run multi-step workflows as containers, with artefact passing and caching |
| Katib | Hyperparameter search and neural architecture search as a Kubernetes job |
| Trainer (formerly Training Operator) | Distributed training (PyTorch, JAX, XGBoost and others) as native Kubernetes objects |
| KServe | Model serving with autoscaling, canaries and scale-to-zero |
| Notebooks | Managed Jupyter servers inside the cluster, near the data and the GPUs |
The current SDK (KFP v2) expresses a step as a decorated Python function. The decorator specifies the base image and any extra packages; KFP builds the container spec for you.
1from kfp import dsl, compiler2from kfp.dsl import Input, Output, Dataset, Model, Metrics34@dsl.component(5 base_image="python:3.11-slim",6 packages_to_install=["pandas==2.2.2", "scikit-learn==1.5.0"],7)8def prepare(source_uri: str, seed: int,9 train_out: Output[Dataset], val_out: Output[Dataset]) -> None:10 import pandas as pd11 from sklearn.model_selection import train_test_split12 df = pd.read_csv(source_uri)13 tr, va = train_test_split(df, test_size=0.2, random_state=seed,14 stratify=df["default"])15 tr.to_parquet(train_out.path)16 va.to_parquet(val_out.path)1718@dsl.component(19 base_image="python:3.11-slim",20 packages_to_install=["pandas==2.2.2", "scikit-learn==1.5.0", "joblib==1.4.2"],21)22def train(train_data: Input[Dataset], val_data: Input[Dataset],23 max_depth: int, n_estimators: int,24 model_out: Output[Model], metrics: Output[Metrics]) -> None:25 import joblib, pandas as pd26 from sklearn.ensemble import RandomForestClassifier27 from sklearn.metrics import roc_auc_score2829 tr, va = pd.read_parquet(train_data.path), pd.read_parquet(val_data.path)30 y_tr, y_va = tr.pop("default"), va.pop("default")3132 m = RandomForestClassifier(max_depth=max_depth, n_estimators=n_estimators,33 random_state=42).fit(tr, y_tr)34 auc = roc_auc_score(y_va, m.predict_proba(va)[:, 1])3536 metrics.log_metric("val_auc", float(auc))37 model_out.metadata["val_auc"] = float(auc)38 joblib.dump(m, model_out.path)3940@dsl.pipeline(name="credit-risk-training",41 description="Prepare then train, with artefact lineage")42def credit_pipeline(source_uri: str, max_depth: int = 6,43 n_estimators: int = 400, seed: int = 42):44 prep = prepare(source_uri=source_uri, seed=seed)45 tr = train(train_data=prep.outputs["train_out"],46 val_data=prep.outputs["val_out"],47 max_depth=max_depth, n_estimators=n_estimators)48 tr.set_cpu_limit("4").set_memory_limit("16Gi")49 # For a GPU step: .set_accelerator_type("nvidia.com/gpu").set_accelerator_limit(1)5051compiler.Compiler().compile(credit_pipeline, "credit_pipeline.yaml")1from kfp.client import Client23client = Client(host="https://kubeflow.internal/pipeline")4run = client.create_run_from_pipeline_package(5 "credit_pipeline.yaml",6 arguments={"source_uri": "s3://ml-data/credit-risk/applications-2024-04-15.csv",7 "max_depth": 8},8 experiment_name="credit-risk",9 enable_caching=True,10)11print(run.run_id)Two features do the heavy lifting. Per-step resources mean the preparation step can ask for 4 CPUs and the training step for a GPU, instead of one machine sized for the worst case and idle the rest of the time. And enable_caching=True means a step whose inputs are unchanged is not re-executed — rerun the pipeline with only max_depth altered and prepare is served from cache.
The cost is real, though. Every step is a container start: image pull, Python startup, package install if you used packages_to_install. That overhead is roughly 30–90 seconds per step. A pipeline of five steps each doing ten seconds of work spends most of its life on scheduling. Kubeflow earns its keep when steps take minutes to hours, not seconds.
Kubeflow earns its keep when a step takes minutes to hours. Below that you are paying container start-up to execute what is really a function call.
Where teams get the boundaries wrong
| Belief | Reality |
|---|---|
| "MLflow versions my data because I logged the CSV as an artefact" | You made a copy attached to one run. Nothing links it to other runs, nothing deduplicates it, and 200 runs means 200 copies. DVC stores one copy per distinct content hash. |
"DVC replaces MLflow — dvc exp show gives me a metrics table" | It does, and it is fine for one person on one machine. It has no shared server, no cross-team search, no model registry and no serving path. |
| "Kubeflow tracks my experiments" | KFP records artefacts and lineage per pipeline run, not a searchable experiment history across ad-hoc local runs. Most Kubeflow users still log to MLflow from inside the steps. |
| "We should adopt all three on day one" | Adopt MLflow first — it gives the largest benefit for the least infrastructure. Add DVC when data versions start causing confusion. Add Kubeflow only when compute is genuinely the constraint and you already run Kubernetes. |
"dvc push is like git push" | They are separate operations on separate stores. Pushing Git without pushing DVC gives a colleague pointer files with no bytes behind them — a confusing "file not found" on a file that visibly exists in the repo. |
The seams: making them work together
Used together, each tool contributes one fact to the same run record. The glue is small:
1import subprocess, mlflow, dvc.api23def dvc_data_hash(path: str) -> str:4 """Read the content hash DVC recorded for this file at the current commit."""5 import yaml6 meta = yaml.safe_load(open(f"{path}.dvc"))7 return meta["outs"][0]["md5"]89with mlflow.start_run():10 mlflow.set_tags({11 "git_sha": subprocess.check_output(["git","rev-parse","HEAD"], text=True).strip(),12 "data_md5": dvc_data_hash("data/raw/applications.csv"),13 "runtime": "kfp", # or "local"14 })15 mlflow.log_params(dvc.api.params_show()["train"])16 ...That single tag block is what turns three tools into one system. Given an MLflow run you can now recover the code (git_sha), the exact data (data_md5 plus dvc pull) and the environment. Any of the three alone leaves a hole.
Choosing, for a team that has none of this
The honest sequencing advice is to install the smallest thing that removes your current pain, and to notice that the pain arrives in a predictable order.
Week one, the pain is "I cannot remember what I tried". Install MLflow locally — it is pip install mlflow and one with block, and it immediately makes the spreadsheet obsolete. Point it at a shared server the moment a second person needs to see your numbers, because local ./mlruns directories are private by definition.
Some weeks later the pain becomes "your numbers and mine disagree and we do not know why". That is the data-versioning pain, and DVC is the cheapest answer. It requires no server at all beyond an S3 bucket you probably already have.
Much later, if ever, the pain becomes "training takes nine hours on the one machine with a GPU and three people are queuing for it". That is when orchestration on shared compute pays for itself. Reaching for Kubeflow before that point means running a Kubernetes cluster to solve a problem you do not have — and the operational cost of that cluster will exceed anything it saves you.
The failure to avoid is the reverse order: standing up a platform first, then discovering nobody logs their runs to it. Tracking is a habit, and habits are easier to build with a tool that takes ten minutes to install than with one that takes a quarter to provision.