MLOps for AI

MLflow, DVC, and Kubeflow Basics


A team of four has been working on a credit-risk model for six weeks. Their record of it is a spreadsheet, experiments_v4.xlsx, with 63 rows and columns headed lr, depth, auc, notes. Row 41 reads 0.912 — best so far!!. Nobody can reproduce it.

The reasons are worth listing because each one points at a different tool:

  • The row does not say which commit produced it, or what the other eleven hyperparameters were set to. Somebody typed four columns and left the rest.
  • It says nothing about the data. Two team members were using data/train_v7.csv; one had regenerated it after a bug fix and kept the same filename. Their AUCs are not comparable and nobody realised.
  • The 0.912 run took nine hours on one person's workstation. It cannot be rerun anywhere else, because "run it" means opening a laptop and executing cells in the right order.

Three failures, three different shapes. One is about recording what happened. One is about versioning large files that cannot go in Git. One is about running work reproducibly on shared compute. MLflow, DVC and Kubeflow each solve one of them well and the other two badly. Teams that treat them as competitors pick one, discover it does not cover everything, and conclude the tool is bad.

QuestionToolWhat it storesWhat it does not do
What did I run, and what came out?MLflowParams, metrics, artefacts, model versions in a databaseDoes not version your data files; does not schedule or provision compute
Which exact bytes of data and model did this use?DVCContent hashes in Git; bytes in object storageNo metric comparison UI worth the name; no serving
Where does this actually execute?KubeflowContainerised pipeline steps on KubernetesNot an experiment tracker; heavy for a single script

These three tools are not alternatives to each other. A team that uses all three uses each for exactly one job, and the seams between them are where the design work lies.

Two different things a spreadsheet was pretending to doMLflow tracks what you tried• Params, metrics and artefacts per run• A run id you can compare against• The model, with its signature• Answers: which configuration won?DVC versions what you fed it• Large files pinned by content hash• A .dvc pointer committed to Git• A pipeline with explicit dependencies• Answers: which data produced that run?
Row 41 was unreproducible because the spreadsheet recorded the score and neither the run nor the data.

MLflow: a database for things you tried

MLflow's core abstraction is the run: one execution of training code, with everything about it recorded in one place. Runs live inside an experiment, which is just a named folder for related runs.

Each run holds four kinds of thing, and knowing which is which prevents most misuse:

CategoryTypeImmutable?ExamplesCommon mistake
ParamsKey → stringYes, set oncemax_depth=6, data_version=v7Trying to log a param twice; logging a metric as a param
MetricsKey → float, with step and timestampNo, append-only seriesval_auc per epochLogging only the final value, losing the curve
TagsKey → stringNo, mutableowner=priya, stage=candidateNot used at all, then search is impossible
ArtefactsFilesYesModel files, plots, confusion matrices, the configLogging a 4 GB dataset as an artefact of every run

The params-are-immutable rule catches people out. If you try to log_param("max_depth", 6) twice with different values in one run, MLflow raises. That is deliberate — a run has one configuration, by definition. If your value changes during training, it is a metric.

Setting up a tracking server

Running with no configuration writes to a local ./mlruns directory, which is fine for the first hour and useless for a team. A shared server needs two backends: a database for the structured records, and object storage for the files.

Bash
pip install "mlflow>=3" psycopg2-binary boto3mlflow server \  --host 0.0.0.0 --port 5000 \  --backend-store-uri postgresql://mlflow:secret@db:5432/mlflow \  --artifacts-destination s3://ml-artifacts/mlflow \  --serve-artifacts

Clients then point at it with an environment variable, never a hardcoded string in the training script:

Bash
export MLFLOW_TRACKING_URI=http://mlflow.internal:5000

Splitting the two stores is the whole reason MLflow scales. Metrics are small and queried constantly, so they belong in Postgres. Model files are large and read rarely, so they belong in object storage. A team that leaves everything on a shared filesystem discovers at about 5,000 runs that listing an experiment takes 40 seconds.

A tracked run, written out

Python
import mlflow, mlflow.sklearnfrom mlflow.models import infer_signaturefrom sklearn.ensemble import RandomForestClassifierfrom sklearn.metrics import roc_auc_score, average_precision_scoremlflow.set_experiment("credit-risk")params = {"n_estimators": 400, "max_depth": 6, "min_samples_leaf": 20,          "class_weight": "balanced", "random_state": 42}with mlflow.start_run(run_name="rf-depth6-balanced") as run:    mlflow.log_params(params)    mlflow.set_tags({"owner": "priya", "data_version": "v7",                     "git_sha": git_sha()})    model = RandomForestClassifier(**params).fit(X_train, y_train)    proba = model.predict_proba(X_val)[:, 1]    mlflow.log_metrics({        "val_auc": roc_auc_score(y_val, proba),        "val_ap":  average_precision_score(y_val, proba),        "val_pos_rate": float(y_val.mean()),    })    # A signature records the expected input/output schema. Without it,    # serving accepts a wrongly-ordered dataframe and silently mispredicts.    signature = infer_signature(X_train, model.predict_proba(X_train))    mlflow.sklearn.log_model(        sk_model=model,        name="model",                 # MLflow 2.x called this `artifact_path`        signature=signature,        input_example=X_train.iloc[:5],    )    mlflow.log_artifact("configs/train_rf.yaml")    print("run_id:", run.info.run_id)

The signature is the line most people omit and later wish they had not. A scikit-learn pipeline fed a DataFrame with the same columns in a different order will not complain; it will use positional order internally and produce confident nonsense. A logged signature makes the serving layer reject that input instead.

Autologging: the fast path and its limits

Python
mlflow.sklearn.autolog()          # or .xgboost, .pytorch, .lightgbm, .transformerswith mlflow.start_run():    model = RandomForestClassifier(n_estimators=400, max_depth=6).fit(X_train, y_train)

That captures every constructor argument as a param, training metrics, the fitted model, and the signature — with no logging code at all. It is genuinely the right default when you are exploring.

What it does not capture is anything the framework does not know about: which data version you used, what your feature engineering did, which git commit is checked out, and any custom metric such as profit at a chosen threshold. So the practical pattern is autolog plus a handful of explicit calls:

Python
mlflow.sklearn.autolog(log_datasets=False)   # datasets can be enormouswith mlflow.start_run():    mlflow.set_tags({"git_sha": git_sha(), "data_version": "v7"})    model = fit(X_train, y_train)    mlflow.log_metric("profit_at_p30", profit_at_threshold(model, X_val, y_val, 0.30))

Wrapping it so the team uses it consistently

Four people logging by hand will produce four naming conventions, and comparison becomes impossible. A thin wrapper enforces the ones that matter.

Python
import contextlib, json, subprocessfrom pathlib import Pathimport mlflowclass RunLogger:    def __init__(self, experiment: str, tracking_uri: str | None = None):        if tracking_uri:            mlflow.set_tracking_uri(tracking_uri)        mlflow.set_experiment(experiment)    @contextlib.contextmanager    def run(self, name: str, config: dict, data_version: str):        dirty = bool(subprocess.check_output(            ["git", "status", "--porcelain"], text=True).strip())        if dirty:            raise RuntimeError("commit before training; a dirty tree makes git_sha a lie")        with mlflow.start_run(run_name=name) as r:            mlflow.log_params(_flatten(config))            mlflow.set_tags({                "git_sha": subprocess.check_output(                    ["git", "rev-parse", "HEAD"], text=True).strip(),                "data_version": data_version,            })            Path("config_snapshot.json").write_text(json.dumps(config, indent=2))            mlflow.log_artifact("config_snapshot.json")            yield rdef _flatten(d: dict, prefix: str = "") -> dict:    """MLflow params are flat; nested config must be flattened first."""    out = {}    for k, v in d.items():        key = f"{prefix}{k}"        if isinstance(v, dict):            out.update(_flatten(v, prefix=f"{key}."))        else:            out[key] = v    return out

_flatten is small and essential. MLflow params are a flat key-value map, so a nested config dictionary logged directly becomes one param whose value is the string "{'model': {'max_depth': 6, ...}}". You cannot filter or sort on that. Flattened, it becomes model.max_depth = 6 and the UI can group by it.

Getting things back out

Python
import mlflowfrom mlflow.tracking import MlflowClient# Query as a DataFrame — the fastest way to compare dozens of runs.runs = mlflow.search_runs(    experiment_names=["credit-risk"],    filter_string="metrics.val_auc > 0.88 and tags.data_version = 'v7'",    order_by=["metrics.val_auc DESC"],    max_results=20,)print(runs[["run_id", "params.model.max_depth", "metrics.val_auc"]])best_id = runs.iloc[0]["run_id"]model = mlflow.pyfunc.load_model(f"runs:/{best_id}/model")preds = model.predict(X_test)# The client gives you the full metric history, not just the last value.client = MlflowClient()for point in client.get_metric_history(best_id, "val_auc"):    print(point.step, round(point.value, 4))

mlflow.pyfunc.load_model is the loader to reach for by default. It returns a uniform predict interface whether the underlying model is scikit-learn, XGBoost or PyTorch, which means downstream code does not have to know or care. Reach for mlflow.sklearn.load_model only when you genuinely need the native object — to inspect feature_importances_, for instance.

DVC: Git for files too big for Git

Git stores every version of every file in the repository. That works beautifully for text and disastrously for a 2.3 GB CSV: the repository grows by 2.3 GB per version, clones take forever, and the delta compression that makes Git efficient does nothing on binary data.

DVC's move is to split the problem. The content goes to object storage. A tiny pointer file containing the content's hash goes into Git. Git then versions the pointers, which is exactly what Git is good at.

Bash
pip install "dvc[s3]"cd credit-riskdvc init                                   # creates .dvc/ and stages it in Gitdvc remote add -d storage s3://ml-data/credit-riskgit commit -m "chore: init dvc"dvc add data/raw/applications.csv          # 2.3 GBgit add data/raw/applications.csv.dvc data/raw/.gitignoregit commit -m "data: applications snapshot 2024-04-15"dvc push                                   # bytes go to S3

The pointer file is about five lines:

Text
outs:- md5: 4b7e2a91c0d83f5e6a1b9c4d7e0f2a38  size: 2451337216  hash: md5  path: applications.csv

A filename is a label; a content hash is an identity. Only one of the two survives somebody quietly regenerating the file.

Now a colleague runs git checkout <commit> followed by dvc pull and gets byte-identical data for that commit. The credit-risk team's second failure — two people silently using different train_v7.csv files — becomes impossible, because the filename is no longer the identity. The hash is.

Pipelines: making the dependency graph explicit

DVC's second capability is a stage graph declared in dvc.yaml. Each stage names its command, its dependencies, its parameters and its outputs.

Text
stages:  prepare:    cmd: python src/prepare.py    deps:      - src/prepare.py      - data/raw/applications.csv    params:      - prepare.test_size      - prepare.seed    outs:      - data/processed/  train:    cmd: python src/train.py    deps:      - src/train.py      - data/processed/    params:      - train.n_estimators      - train.max_depth    outs:      - models/model.joblib    metrics:      - reports/metrics.json:          cache: false    plots:      - reports/roc.csv:          x: fpr          y: tpr

With params.yaml alongside it:

Text
prepare:  test_size: 0.2  seed: 42train:  n_estimators: 400  max_depth: 6

Then dvc repro hashes every dependency and every parameter, compares against the last run, and re-executes only the stages whose inputs changed. Change max_depth from 6 to 8 and prepare is skipped entirely — a real saving when preparation takes forty minutes. Change the raw data and both stages rerun.

Bash
dvc repro                        # rebuild what is staledvc metrics diff HEAD~1          # what did that change do to the numbers?dvc exp run -S train.max_depth=8 # a throwaway experiment, no commit neededdvc exp show                     # table of experiments with params and metrics

DVC's Python API lets other code read versioned data without checking it out:

Python
import dvc.apiimport pandas as pd# Read a specific historical version straight from remote storage.with dvc.api.open("data/raw/applications.csv",                  repo="https://github.com/acme/credit-risk",                  rev="v1.2.0", mode="r") as f:    df = pd.read_csv(f)params = dvc.api.params_show()           # the params.yaml for the current commiturl = dvc.api.get_url("models/model.joblib", rev="v1.2.0")

Kubeflow: running the work somewhere that is not a laptop

The third failure — a nine-hour run that only exists on one workstation — is a compute problem. Kubeflow is a set of Kubernetes components for machine learning; the relevant piece here is Kubeflow Pipelines, which turns each step of a workflow into a container that Kubernetes schedules.

ComponentPurpose
Pipelines (KFP)Define and run multi-step workflows as containers, with artefact passing and caching
KatibHyperparameter search and neural architecture search as a Kubernetes job
Trainer (formerly Training Operator)Distributed training (PyTorch, JAX, XGBoost and others) as native Kubernetes objects
KServeModel serving with autoscaling, canaries and scale-to-zero
NotebooksManaged Jupyter servers inside the cluster, near the data and the GPUs

The current SDK (KFP v2) expresses a step as a decorated Python function. The decorator specifies the base image and any extra packages; KFP builds the container spec for you.

Python
from kfp import dsl, compilerfrom kfp.dsl import Input, Output, Dataset, Model, Metrics@dsl.component(    base_image="python:3.11-slim",    packages_to_install=["pandas==2.2.2", "scikit-learn==1.5.0"],)def prepare(source_uri: str, seed: int,            train_out: Output[Dataset], val_out: Output[Dataset]) -> None:    import pandas as pd    from sklearn.model_selection import train_test_split    df = pd.read_csv(source_uri)    tr, va = train_test_split(df, test_size=0.2, random_state=seed,                              stratify=df["default"])    tr.to_parquet(train_out.path)    va.to_parquet(val_out.path)@dsl.component(    base_image="python:3.11-slim",    packages_to_install=["pandas==2.2.2", "scikit-learn==1.5.0", "joblib==1.4.2"],)def train(train_data: Input[Dataset], val_data: Input[Dataset],          max_depth: int, n_estimators: int,          model_out: Output[Model], metrics: Output[Metrics]) -> None:    import joblib, pandas as pd    from sklearn.ensemble import RandomForestClassifier    from sklearn.metrics import roc_auc_score    tr, va = pd.read_parquet(train_data.path), pd.read_parquet(val_data.path)    y_tr, y_va = tr.pop("default"), va.pop("default")    m = RandomForestClassifier(max_depth=max_depth, n_estimators=n_estimators,                               random_state=42).fit(tr, y_tr)    auc = roc_auc_score(y_va, m.predict_proba(va)[:, 1])    metrics.log_metric("val_auc", float(auc))    model_out.metadata["val_auc"] = float(auc)    joblib.dump(m, model_out.path)@dsl.pipeline(name="credit-risk-training",              description="Prepare then train, with artefact lineage")def credit_pipeline(source_uri: str, max_depth: int = 6,                    n_estimators: int = 400, seed: int = 42):    prep = prepare(source_uri=source_uri, seed=seed)    tr = train(train_data=prep.outputs["train_out"],               val_data=prep.outputs["val_out"],               max_depth=max_depth, n_estimators=n_estimators)    tr.set_cpu_limit("4").set_memory_limit("16Gi")    # For a GPU step: .set_accelerator_type("nvidia.com/gpu").set_accelerator_limit(1)compiler.Compiler().compile(credit_pipeline, "credit_pipeline.yaml")
Python
from kfp.client import Clientclient = Client(host="https://kubeflow.internal/pipeline")run = client.create_run_from_pipeline_package(    "credit_pipeline.yaml",    arguments={"source_uri": "s3://ml-data/credit-risk/applications-2024-04-15.csv",               "max_depth": 8},    experiment_name="credit-risk",    enable_caching=True,)print(run.run_id)

Two features do the heavy lifting. Per-step resources mean the preparation step can ask for 4 CPUs and the training step for a GPU, instead of one machine sized for the worst case and idle the rest of the time. And enable_caching=True means a step whose inputs are unchanged is not re-executed — rerun the pipeline with only max_depth altered and prepare is served from cache.

The cost is real, though. Every step is a container start: image pull, Python startup, package install if you used packages_to_install. That overhead is roughly 30–90 seconds per step. A pipeline of five steps each doing ten seconds of work spends most of its life on scheduling. Kubeflow earns its keep when steps take minutes to hours, not seconds.

Kubeflow earns its keep when a step takes minutes to hours. Below that you are paying container start-up to execute what is really a function call.

Where teams get the boundaries wrong

BeliefReality
"MLflow versions my data because I logged the CSV as an artefact"You made a copy attached to one run. Nothing links it to other runs, nothing deduplicates it, and 200 runs means 200 copies. DVC stores one copy per distinct content hash.
"DVC replaces MLflow — dvc exp show gives me a metrics table"It does, and it is fine for one person on one machine. It has no shared server, no cross-team search, no model registry and no serving path.
"Kubeflow tracks my experiments"KFP records artefacts and lineage per pipeline run, not a searchable experiment history across ad-hoc local runs. Most Kubeflow users still log to MLflow from inside the steps.
"We should adopt all three on day one"Adopt MLflow first — it gives the largest benefit for the least infrastructure. Add DVC when data versions start causing confusion. Add Kubeflow only when compute is genuinely the constraint and you already run Kubernetes.
"dvc push is like git push"They are separate operations on separate stores. Pushing Git without pushing DVC gives a colleague pointer files with no bytes behind them — a confusing "file not found" on a file that visibly exists in the repo.

The seams: making them work together

Used together, each tool contributes one fact to the same run record. The glue is small:

Python
import subprocess, mlflow, dvc.apidef dvc_data_hash(path: str) -> str:    """Read the content hash DVC recorded for this file at the current commit."""    import yaml    meta = yaml.safe_load(open(f"{path}.dvc"))    return meta["outs"][0]["md5"]with mlflow.start_run():    mlflow.set_tags({        "git_sha":  subprocess.check_output(["git","rev-parse","HEAD"], text=True).strip(),        "data_md5": dvc_data_hash("data/raw/applications.csv"),        "runtime":  "kfp",     # or "local"    })    mlflow.log_params(dvc.api.params_show()["train"])    ...

That single tag block is what turns three tools into one system. Given an MLflow run you can now recover the code (git_sha), the exact data (data_md5 plus dvc pull) and the environment. Any of the three alone leaves a hole.

Choosing, for a team that has none of this

The honest sequencing advice is to install the smallest thing that removes your current pain, and to notice that the pain arrives in a predictable order.

Week one, the pain is "I cannot remember what I tried". Install MLflow locally — it is pip install mlflow and one with block, and it immediately makes the spreadsheet obsolete. Point it at a shared server the moment a second person needs to see your numbers, because local ./mlruns directories are private by definition.

Some weeks later the pain becomes "your numbers and mine disagree and we do not know why". That is the data-versioning pain, and DVC is the cheapest answer. It requires no server at all beyond an S3 bucket you probably already have.

Much later, if ever, the pain becomes "training takes nine hours on the one machine with a GPU and three people are queuing for it". That is when orchestration on shared compute pays for itself. Reaching for Kubeflow before that point means running a Kubernetes cluster to solve a problem you do not have — and the operational cost of that cluster will exceed anything it saves you.

The failure to avoid is the reverse order: standing up a platform first, then discovering nobody logs their runs to it. Tracking is a habit, and habits are easier to build with a tool that takes ten minutes to install than with one that takes a quarter to provision.