MLOps for AI

AWS SageMaker, Vertex AI, and Hugging Face Hub


Five data scientists share one GPU box. It is a g5.2xlarge, on-demand at about USD 1.21 an hour (the us-east-1 list price at the time of writing), and it runs continuously because nobody wants to be the person who terminates an instance somebody else is using. Over a 30-day month that is 720 × 1.212 = USD 872.64.

The machine is actually busy about 30% of the time — roughly 216 hours. So the true cost of a useful GPU-hour is 872.64 / 216 = USD 4.04, more than three times the sticker price. Worse than the money: because there is one box, experiments queue. A six-hour training run started at 16:00 means nobody else trains until 22:00, so the team's experiment cycle is measured in days.

Then the deployment problem arrives. The model has to be served, with TLS, autoscaling, health checks, a rollback path and an audit trail. Building that is four to six weeks of engineering for a team of five whose job is modelling.

Managed ML platforms exist to sell you out of both problems. It is worth being precise about what you are buying, because the pitch ("end-to-end ML platform") obscures the specifics.

A managed endpoint and a hub are not competitorsManaged — SageMaker, Vertex AI• Autoscaling, traffic splits, monitoring• Billed per instance-hour, idle included• A model package group gates promotion• Leaving costs you a rewriteHub — Hugging Face• A versioned repo plus a model card• Free to host, you pay to serve• Inference Endpoints when you need an SLA• Runs anywhere transformers runs
The 872 USD month was not a platform choice — it was an on-demand box nobody felt entitled to terminate.

What you are actually buying

CapabilityWhat it replacesRealistic value
Ephemeral training jobsAn always-on GPU instanceBilled per second while running. The 216 busy hours cost 216 × 1.515 = USD 327 even at the platform's ~25% premium — under 40% of the idle-box bill
Managed spot / preemptiblePaying on-demand for interruptible work60–90% off, if your job checkpoints and can resume
Managed endpointsLoad balancer, autoscaler, TLS, health checks, blue-greenWeeks of platform engineering you do not do
Built-in registry and lineageA registry you run yourselfModerate — good enough, and tied to the platform
Traffic splittingA custom routerHigh. Canary and A/B become configuration
Drift monitoringA monitoring service you buildGenuinely useful, and the least portable thing you will adopt

You are not buying machine learning. You are buying the operational plumbing around it — and paying a margin plus a portability cost for the privilege.

The three platforms below are not equivalent. SageMaker and Vertex AI are full lifecycle platforms tied to one cloud. The Hugging Face Hub is a model distribution platform that also happens to offer hosting, and it plays a different role: it is where models come from and where open models go, whichever cloud you deploy on.

AWS SageMaker

A training job as a Python object

The unit of work is an estimator: your script, a container image, an instance type, and pointers to S3.

The code in this section uses the SageMaker Python SDK v2 (pip install "sagemaker<3"), which is what most existing pipelines run. SDK v3 replaced the Estimator, Model and Predictor classes with ModelTrainer and ModelBuilder and has no compatibility layer, so check which major version a project pins before copying code between them. The platform concepts — jobs, channels, model packages, endpoints, variants — are the same in both.

Python
import sagemakerfrom sagemaker.sklearn.estimator import SKLearnsession = sagemaker.Session()role = "arn:aws:iam::123456789012:role/SageMakerExecutionRole"bucket = session.default_bucket()estimator = SKLearn(    entry_point="train.py",    source_dir="src/",                 # requirements.txt here is installed for you    framework_version="1.2-1",    py_version="py3",    role=role,    instance_type="ml.m5.2xlarge",    instance_count=1,    hyperparameters={"n-estimators": 400, "max-depth": 6, "seed": 42},    metric_definitions=[               # scraped from stdout into CloudWatch        {"Name": "val:auc", "Regex": r"val_auc=([0-9\.]+)"},    ],    output_path=f"s3://{bucket}/churn/models",    checkpoint_s3_uri=f"s3://{bucket}/churn/checkpoints",    use_spot_instances=True,    max_run=3600,    max_wait=7200,                     # must exceed max_run when using spot    environment={"MLFLOW_TRACKING_URI": "http://mlflow.internal:5000"},)estimator.fit({    "train": f"s3://{bucket}/churn/data/train/",    "validation": f"s3://{bucket}/churn/data/validation/",}, job_name="churn-gbm-2024-04-15-01")

Four details carry most of the value. use_spot_instances with a checkpoint_s3_uri typically cuts the bill by 60–70%; without checkpoints, an interruption at hour five of a six-hour job costs you the whole job. max_wait must be larger than max_run or the job fails validation, because the wait covers queueing for spot capacity as well as running. metric_definitions turns printed lines into real time-series — your script simply prints val_auc=0.8871 and the platform harvests it. And source_dir means a requirements.txt beside your script gets installed, so you rarely need a custom image.

Inside train.py, the platform communicates through environment variables and fixed paths:

Python
import argparse, os, joblibimport pandas as pdp = argparse.ArgumentParser()p.add_argument("--n-estimators", type=int, default=400)p.add_argument("--max-depth", type=int, default=6)p.add_argument("--seed", type=int, default=42)p.add_argument("--train", default=os.environ["SM_CHANNEL_TRAIN"])p.add_argument("--validation", default=os.environ["SM_CHANNEL_VALIDATION"])p.add_argument("--model-dir", default=os.environ["SM_MODEL_DIR"])args = p.parse_args()train = pd.read_parquet(args.train)# ... fit ...print(f"val_auc={auc:.4f}")            # harvested by metric_definitionsjoblib.dump(model, os.path.join(args.model_dir, "model.joblib"))

Anything written to SM_MODEL_DIR is tarred and uploaded to S3 when the job ends. Anything written anywhere else is destroyed with the container — a mistake people make exactly once.

On a managed platform the container is the deliverable, and everything written outside the designated output directory is destroyed when the job ends.

Training a Hugging Face model on SageMaker

Python
from sagemaker.huggingface import HuggingFacehf = HuggingFace(    entry_point="train_transformer.py",    source_dir="src/",    instance_type="ml.g5.2xlarge",    instance_count=1,    role=role,    transformers_version="4.36",    pytorch_version="2.1",    py_version="py310",    hyperparameters={        "model_name_or_path": "distilbert-base-uncased",        "epochs": 3, "per_device_train_batch_size": 32, "learning_rate": 5e-5,    },    use_spot_instances=True, max_run=7200, max_wait=14400,)hf.fit({"train": f"s3://{bucket}/reviews/train", "test": f"s3://{bucket}/reviews/test"})

The point of the dedicated estimator is that the deep-learning container is already built with matching CUDA, PyTorch and transformers versions. Assembling that stack yourself is a genuinely unpleasant afternoon.

Pipelines and the model package group

A single job is not a system. SageMaker Pipelines chains steps into a DAG whose parameters can be overridden per execution.

Python
from sagemaker.workflow.pipeline import Pipelinefrom sagemaker.workflow.parameters import ParameterString, ParameterFloatfrom sagemaker.workflow.steps import TrainingStep, ProcessingStepfrom sagemaker.workflow.condition_step import ConditionStepfrom sagemaker.workflow.conditions import ConditionGreaterThanOrEqualTofrom sagemaker.workflow.model_step import ModelStepfrom sagemaker.workflow.functions import JsonGetdata_uri  = ParameterString("DataUri", default_value=f"s3://{bucket}/churn/data/")min_auc   = ParameterFloat("MinAuc", default_value=0.85)train_step = TrainingStep(name="Train", estimator=estimator,                          inputs={"train": data_uri})eval_step  = ProcessingStep(name="Evaluate", processor=sklearn_processor,                            code="src/evaluate.py", property_files=[eval_report])gate = ConditionStep(    name="AucGate",    conditions=[ConditionGreaterThanOrEqualTo(        left=JsonGet(step_name="Evaluate", property_file=eval_report,                     json_path="metrics.val_auc"),        right=min_auc)],    if_steps=[ModelStep(name="Register", step_args=model.register(        content_types=["application/json"], response_types=["application/json"],        inference_instances=["ml.m5.large"],        model_package_group_name="churn-models",        approval_status="PendingManualApproval"))],    else_steps=[],)Pipeline(name="churn-training",         parameters=[data_uri, min_auc],         steps=[train_step, eval_step, gate]).upsert(role_arn=role)

approval_status="PendingManualApproval" is the important choice. A model package that trains successfully is registered, not deployable. Someone — or an automated check — must approve it, and that approval is recorded. Setting this to Approved because the manual step is annoying converts your pipeline into an unattended way to deploy whatever came out of training.

Endpoints, traffic splits and autoscaling

A SageMaker endpoint hosts one or more production variants, each with its own model, instance type and traffic weight.

Python
import boto3sm = boto3.client("sagemaker")sm.create_endpoint_config(    EndpointConfigName="churn-config-v4",    ProductionVariants=[        {"VariantName": "champion", "ModelName": "churn-v13",         "InstanceType": "ml.m5.large", "InitialInstanceCount": 2,         "InitialVariantWeight": 9.0},        {"VariantName": "challenger", "ModelName": "churn-v14",         "InstanceType": "ml.m5.large", "InitialInstanceCount": 1,         "InitialVariantWeight": 1.0},    ],    DataCaptureConfig={                       # required for later drift analysis        "EnableCapture": True,        "InitialSamplingPercentage": 20,        "DestinationS3Uri": f"s3://{bucket}/churn/capture",        "CaptureOptions": [{"CaptureMode": "Input"}, {"CaptureMode": "Output"}],    },)sm.update_endpoint(EndpointName="churn", EndpointConfigName="churn-config-v4")# Shift traffic with no downtime and no redeploy.sm.update_endpoint_weights_and_capacities(    EndpointName="churn",    DesiredWeightsAndCapacities=[        {"VariantName": "champion",  "DesiredWeight": 5.0},        {"VariantName": "challenger","DesiredWeight": 5.0},    ],)

Weights are relative, not percentages. With champion=9 and challenger=1 the challenger receives 1/(9+1) = 10% of requests. People assume the numbers are percentages, set 90 and 10, and get exactly the same split — which works by coincidence and stops working the moment a third variant appears. Rolling back is one call with the challenger weight set to 0.

Autoscaling is configured through Application Auto Scaling rather than SageMaker itself:

Python
aas = boto3.client("application-autoscaling")resource_id = "endpoint/churn/variant/champion"aas.register_scalable_target(    ServiceNamespace="sagemaker", ResourceId=resource_id,    ScalableDimension="sagemaker:variant:DesiredInstanceCount",    MinCapacity=2, MaxCapacity=12,)aas.put_scaling_policy(    PolicyName="churn-invocations", ServiceNamespace="sagemaker",    ResourceId=resource_id,    ScalableDimension="sagemaker:variant:DesiredInstanceCount",    PolicyType="TargetTrackingScaling",    TargetTrackingScalingPolicyConfiguration={        "TargetValue": 1200.0,       # invocations per instance per minute        "PredefinedMetricSpecification": {            "PredefinedMetricType": "SageMakerVariantInvocationsPerInstance"},        "ScaleInCooldown": 600, "ScaleOutCooldown": 60,    },)

The asymmetric cooldowns are deliberate: scale out fast when load arrives (60 s), scale in slowly (600 s) so a brief lull does not remove capacity you need again two minutes later. MinCapacity=2, not 1, because a single instance means a single availability zone and no redundancy.

Google Vertex AI

Vertex covers the same ground with a different shape: models are versioned resources in a registry, endpoints hold deployed models, and traffic split is a dictionary that must sum to 100.

Python
from google.cloud import aiplatformaiplatform.init(project="acme-ml", location="europe-west4",                staging_bucket="gs://acme-ml-staging")# Upload as a NEW VERSION of an existing model, not a new model.model = aiplatform.Model.upload(    parent_model="projects/acme-ml/locations/europe-west4/models/churn",    display_name="churn",    version_aliases=["challenger"],    artifact_uri="gs://acme-ml/churn/v14/",    serving_container_image_uri=(        "europe-docker.pkg.dev/vertex-ai/prediction/sklearn-cpu.1-3:latest"),    version_description="Adds tenure_bucket feature; val_auc 0.8912",    labels={"team": "growth", "git_sha": "6f2a91c"},)

Omitting parent_model is the classic Vertex mistake: you get a brand-new model resource rather than version 14 of the existing one, and your version history quietly splits into two lineages that nothing connects.

Python
endpoint = aiplatform.Endpoint("projects/acme-ml/locations/europe-west4/endpoints/123")endpoint.deploy(    model=model,    deployed_model_display_name="churn-v14",    machine_type="n1-standard-4",    min_replica_count=2, max_replica_count=10,    traffic_split={"1234567890": 90, "0": 10},   # "0" = the model deployed by this call    autoscaling_target_cpu_utilization=60,)# Shift or roll back without redeploying anything.endpoint.update(traffic_split={"1234567890": 0, "5678901234": 100})

Unlike SageMaker's relative weights, Vertex percentages must sum to exactly 100 or the call is rejected. The special key "0" refers to the model being deployed by this very call, which makes the initial split expressible before the new deployed-model ID exists; the other key is the incumbent's deployed-model ID (1234567890 here). Once the deploy finishes, the new model has its own ID (5678901234), and that is what later update calls use.

Pipelines and monitoring

Vertex Pipelines runs KFP-compiled pipelines as a managed service, so the pipeline definition is ordinary Python with decorated components:

Python
from kfp import dsl, compilerfrom google.cloud import aiplatform@dsl.component(base_image="python:3.11",               packages_to_install=["pandas==2.2.2", "scikit-learn==1.5.0"])def train(data_uri: str, max_depth: int, model_out: dsl.Output[dsl.Model]):    ...@dsl.pipeline(name="churn-training", pipeline_root="gs://acme-ml/pipeline-root")def churn(data_uri: str, max_depth: int = 6):    train(data_uri=data_uri, max_depth=max_depth)compiler.Compiler().compile(churn, "churn.yaml")aiplatform.PipelineJob(    display_name="churn-training",    template_path="churn.yaml",    parameter_values={"data_uri": "gs://acme-ml/churn/2024-04-15/"},    enable_caching=True,).submit(service_account="ml-pipelines@acme-ml.iam.gserviceaccount.com")

Model monitoring compares live request features against a training baseline and alerts on drift:

Python
from google.cloud.aiplatform import model_monitoringaiplatform.ModelDeploymentMonitoringJob.create(    display_name="churn-monitoring",    endpoint=endpoint,    logging_sampling_strategy=model_monitoring.RandomSampleConfig(sample_rate=0.2),    schedule_config=model_monitoring.ScheduleConfig(monitor_interval=1),  # hours    alert_config=model_monitoring.EmailAlertConfig(        user_emails=["ml-oncall@acme.com"], enable_logging=True),    objective_configs=model_monitoring.ObjectiveConfig(        skew_detection_config=model_monitoring.SkewDetectionConfig(            data_source="bq://acme-ml.churn.training_snapshot",            target_field="churned",            skew_thresholds={"tenure_months": 0.15, "monthly_charges": 0.15},        ),        drift_detection_config=model_monitoring.DriftDetectionConfig(            drift_thresholds={"tenure_months": 0.15, "monthly_charges": 0.15},        ),    ),)

The two configs answer different questions and both are worth having. Skew compares live traffic against the training data — it catches a training/serving mismatch that has been there since day one. Drift compares live traffic against recent live traffic — it catches the world changing after deployment. A team that configures only drift will never discover that their serving preprocessing was subtly different from training all along, because the comparison never looks back that far.

Hugging Face Hub

The Hub is a different kind of thing: a version-controlled repository host for models, datasets and demos, backed by Git, with large files held in Hugging Face's Xet storage (the successor to Git LFS, which still works for pushing). Its value is distribution and provenance rather than infrastructure.

Bash
pip install "huggingface_hub>=1.0" transformershf auth login                  # or set HF_TOKEN in the environment
Python
from huggingface_hub import HfApi, create_repo, upload_folderapi = HfApi()create_repo("acme/review-sentiment", repo_type="model",            private=True, exist_ok=True)upload_folder(    repo_id="acme/review-sentiment",    folder_path="artifacts/distilbert-sentiment/",    commit_message="v2: 3 epochs, lr 5e-5, macro-F1 0.913",    ignore_patterns=["*.ipynb", "checkpoints/*", "*.log"],)# Tag the commit so it can be pinned by consumers.api.create_tag("acme/review-sentiment", tag="v2.0.0",               tag_message="Production release 2024-04-15")

For a transformers model the one-liner is model.push_to_hub("acme/review-sentiment") alongside tokenizer.push_to_hub(...). Forgetting the tokenizer is the single most common Hub mistake: the model loads, tokenises with a default vocabulary, and produces confident nonsense.

The consumer side is where pinning matters:

Python
from transformers import AutoModelForSequenceClassification, AutoTokenizerREV = "v2.0.0"      # a tag or commit SHA - never the default "main"model = AutoModelForSequenceClassification.from_pretrained(    "acme/review-sentiment", revision=REV)tokenizer = AutoTokenizer.from_pretrained("acme/review-sentiment", revision=REV)

Loading without revision means "whatever is on main right now". A colleague pushing an experiment changes your production behaviour with no deploy on your side.

The model card is documentation with a schema

Python
from huggingface_hub import ModelCard, ModelCardDatacard = ModelCard.from_template(    card_data=ModelCardData(        language="en", license="apache-2.0",        library_name="transformers", tags=["sentiment", "distilbert"],        datasets=["acme/product-reviews"],        metrics=["f1", "accuracy"],        model_name="review-sentiment",    ),    model_description=(        "DistilBERT fine-tuned on 240k English product reviews to classify "        "sentiment as negative / neutral / positive."    ),    intended_uses=(        "Triaging incoming product reviews for the support queue. "        "NOT validated for medical, legal or financial text."    ),    limitations=(        "Macro-F1 is 0.913 overall but 0.71 on reviews shorter than 8 tokens. "        "English only; code-switched text degrades sharply. "        "Trained on reviews up to 2024-02; newer product names are unseen."    ),    training_data="acme/product-reviews, 240,118 rows, split 80/10/10 by review_id.",    evaluation="Held-out 24,012 reviews. Macro-F1 0.913, accuracy 0.928.",)card.push_to_hub("acme/review-sentiment")

The limitations section is the part with real operational value. "Macro-F1 0.71 on reviews shorter than 8 tokens" is the sentence that stops a downstream team routing one-word reviews through this model and wondering why the triage queue is wrong.

From widget to managed endpoint

The inference widget on a model page is a demo: shared capacity, no SLA, and for most private or custom models no shared provider serves them at all. Production use needs a dedicated endpoint.

Python
from huggingface_hub import create_inference_endpointep = create_inference_endpoint(    name="review-sentiment-prod",    repository="acme/review-sentiment",    revision="v2.0.0",    framework="pytorch", task="text-classification",    accelerator="gpu", instance_size="x1", instance_type="nvidia-t4",    vendor="aws", region="eu-west-1", type="authenticated",    min_replica=1, max_replica=4,)ep.wait()print(ep.url, ep.status)

type="authenticated" (called "protected" in older versions) requires a Hugging Face token on every request; "public" means anyone with the URL can invoke it and you pay for their traffic. Setting min_replica=0 together with a scale_to_zero_timeout enables scale-to-zero, which removes idle cost at the price of a cold start of roughly 30–60 seconds on the first request after a quiet period — fine for internal batch tooling, unacceptable behind a user-facing form.

Comparing the three

DimensionSageMakerVertex AIHugging Face Hub
Primary roleFull lifecycle on AWSFull lifecycle on GCPModel distribution, provenance, light hosting
TrainingEstimators, spot, distributed, PipelinesCustom jobs, Vertex Pipelines (KFP)None — you train elsewhere
RegistryModel package groups with approval statusModel resource with versions and aliasesGit repos with tags and commit SHAs
Traffic splittingRelative variant weightsPercentages summing to 100Not really; separate endpoints
AutoscalingApplication Auto Scaling, target trackingBuilt into deploy, CPU-target basedReplica range, optional scale-to-zero
Drift monitoringModel Monitor with data captureBuilt-in skew and drift jobsNone
Lock-inHigh — SDK, IAM, S3, container conventionsHigh — SDK, IAM, GCSLow — standard files in a Git repo
Cost shapePer-second compute plus a hosting premiumPer-node-hourFree public hosting; per-hour for endpoints
Best atDeep AWS integration, mature spot trainingPipelines and monitoring feel most coherentSharing, discovery, open models, model cards

How to choose without agonising

The decision is dominated by one question that has nothing to do with machine learning: where does your data already live? Training reads terabytes. Egress across clouds is slow and metered, and the cost of moving data typically dwarfs any feature difference between platforms. If your warehouse is BigQuery, use Vertex. If it is S3 and Redshift, use SageMaker.

Data gravity beats feature comparison. The platform your data already sits next to wins almost every argument about which one is technically better.

SituationChooseWhy
Data in S3, team knows AWSSageMakerNo egress, existing IAM, richest spot support
Data in BigQueryVertex AINative BigQuery integration; the pipeline story is cleaner
Publishing an open modelHugging Face HubIt is where people look, with cards and versioning built in
Fine-tuning an open LLMHub for weights, cloud for computePull from the Hub, train on SageMaker or Vertex, push results back
Multi-cloud requirementKubernetes plus the HubManaged platforms are the least portable layer you can adopt
Two people, one model, 300 requests/dayNone of themA container on a small VM behind a load balancer costs 20 dollars a month and takes an afternoon

The failure modes that cost real money

Endpoints left running. This is the classic and it is expensive. A training job stops billing when it finishes; an endpoint bills until you delete it. An ml.g5.2xlarge endpoint left up after a demo costs roughly USD 1.5 per hour at current list prices, which is USD 1,080 over a month nobody looked at it. Tag every endpoint with an owner and an expiry, and run a scheduled job that deletes untagged ones.

Spot without checkpoints. Enabling use_spot_instances and omitting checkpoint_s3_uri converts a 70% discount into a lottery. An interruption at hour five of six loses everything and you rerun at full length.

Assuming weights are percentages. SageMaker variant weights are relative. Two variants at 90 and 10 behave identically to 9 and 1; add a third at 10 and the arithmetic changes under you.

Loading from main. An unpinned Hub revision means someone else's push changes your production behaviour. Pin a tag or a commit SHA, always.

Building the platform into the model code. The most expensive mistake is not operational but architectural: scattering sagemaker. and aiplatform. imports through your training and feature code. Keep the platform at the edges — your train.py should read paths from arguments and write to a directory, knowing nothing about who invoked it. That way the same script runs locally, in a container, on SageMaker and on Vertex, and the day your company signs a deal with a different cloud you rewrite a launcher, not a codebase.