Model Deployment for AI Engineers

Cloud Deployment (AWS SageMaker, GCP Vertex AI, Hugging Face Spaces, AWS Lambda)


A startup deploys an image classifier to a SageMaker real-time endpoint on an ml.g4dn.xlarge GPU instance. It works beautifully. Six weeks later finance flags a cloud bill of roughly 540 USD for that one endpoint. The traffic logs show 8,400 predictions that month — about twelve per hour. They paid for 720 hours of an idle GPU to serve twelve requests an hour.

The same workload on AWS Lambda would have cost under 1 USD. The model was fine, the code was fine, the platform was fine — the match between traffic shape and pricing model was catastrophically wrong.

That is the recurring theme of cloud deployment. These platforms are not ranked best to worst. They are priced and engineered for different traffic patterns, and picking the wrong one costs you either an order of magnitude in money or an order of magnitude in latency. This lesson is about telling them apart.

Where the always-on endpoint starts paying for itself4.02 USD165.60 USDLambda40.20 USD165.60 USDLambda165.60 USD165.60 USDBreak-even804 USD165.60 USDEndpointLambda, 3 GBm5.xlarge endpointCheaper100,000 / month1,000,000 / month4,120,000 / month20,000,000 / monthAt twelve predictions an hour the 540 USD GPU endpoint sat idle 99.9 percent of the month at 6.4 cents a call.
An always-on endpoint is a fixed cost pretending to be a per-request one, and the crossover is higher than it feels.

Four different things called "deployment"

PatternYou pay forLatencyFits
Managed endpoint SageMaker, Vertex AIInstance-hours, whether or not requests arriveConsistently 50–200 msSteady traffic above roughly 2 req/s
Serverless function Lambda, Cloud RunInvocations and GB-seconds only50 ms warm, 3–15 s coldSpiky, low, or unpredictable traffic
Batch / async job Batch Transform, Batch PredictionCompute for the duration of the jobMinutes to hoursScoring millions of rows overnight
Hosted demo Hugging Face SpacesNothing, on the free tierSeconds, sleeps when idlePrototypes, stakeholder review, public demos

Before choosing anything, answer three questions with actual numbers: how many requests per second at peak, how many per month in total, and what latency does the user experience require? Those three numbers determine the answer almost entirely. Teams that skip them tend to reach for whichever platform they read about most recently.

AWS SageMaker: managed endpoints

The mental model: you hand SageMaker a container image and a model artefact in S3. It runs the container on instances it manages, puts a load balancer and TLS in front, and gives you an HTTPS endpoint. The container must implement two routes — GET /ping for health and POST /invocations for inference — and listen on port 8080. That is the entire contract, which means any FastAPI server you already have becomes a SageMaker container with almost no change.

The deployment code below uses the SageMaker Python SDK v2 (pip install "sagemaker<3"). SDK v3 replaced the PyTorchModel-style classes with a single ModelBuilder and does not run v2 code unchanged, so check which major version your project pins. The container contract and the four inference hooks are the same either way.

Python
import boto3, sagemakerfrom sagemaker.pytorch import PyTorchModelrole = "arn:aws:iam::123456789012:role/SageMakerExecutionRole"model = PyTorchModel(    model_data="s3://my-bucket/models/defectnet/2.5.0/model.tar.gz",    role=role,    entry_point="inference.py",     # defines the four handler functions below    source_dir="code/",    framework_version="2.3",    py_version="py311",    env={"OMP_NUM_THREADS": "1", "MODEL_VERSION": "2.5.0"},)predictor = model.deploy(    initial_instance_count=2,       # two for availability, not for throughput    instance_type="ml.m5.xlarge",    endpoint_name="defect-classifier-prod",)

The artefact must be a .tar.gz, and SageMaker extracts it into /opt/ml/model inside the container. Your inference.py defines four hooks:

Python
import torch, json, iofrom PIL import Imageimport torchvision.transforms as Tdef model_fn(model_dir):                         # called ONCE at container start    m = torch.jit.load(f"{model_dir}/model.pt", map_location="cpu")    m.eval()    return mdef input_fn(request_body, content_type):        # deserialise    if content_type == "application/x-image":        img = Image.open(io.BytesIO(request_body)).convert("RGB")        tfm = T.Compose([T.Resize((224, 224)), T.ToTensor(),                         T.Normalize([0.485, 0.456, 0.406], [0.229, 0.224, 0.225])])        return tfm(img).unsqueeze(0)    raise ValueError(f"unsupported content type: {content_type}")def predict_fn(x, model):                        # inference    with torch.no_grad():        return torch.softmax(model(x), dim=1)def output_fn(probs, accept):                    # serialise    idx = int(probs.argmax())    return json.dumps({"label": LABELS[idx],                       "confidence": round(float(probs[0, idx]), 4)}), "application/json"

Splitting the pipeline into these four functions is not arbitrary. It lets SageMaker call model_fn exactly once per container while calling the other three per request — which is the same "load at startup, not in the handler" rule that governs every serving stack, enforced by the framework.

Auto-scaling

A fixed instance count is either wasteful at night or insufficient at peak. Target tracking adjusts the count to hold a chosen metric at a chosen value.

Python
aas = boto3.client("application-autoscaling")resource_id = "endpoint/defect-classifier-prod/variant/AllTraffic"aas.register_scalable_target(    ServiceNamespace="sagemaker", ResourceId=resource_id,    ScalableDimension="sagemaker:variant:DesiredInstanceCount",    MinCapacity=2, MaxCapacity=12,)aas.put_scaling_policy(    PolicyName="invocations-per-instance",    ServiceNamespace="sagemaker", ResourceId=resource_id,    ScalableDimension="sagemaker:variant:DesiredInstanceCount",    PolicyType="TargetTrackingScaling",    TargetTrackingScalingPolicyConfiguration={        "TargetValue": 1200.0,       # invocations per instance per minute        "PredefinedMetricSpecification": {            "PredefinedMetricType": "SageMakerVariantInvocationsPerInstance"        },        "ScaleInCooldown": 600,      # slow to shrink        "ScaleOutCooldown": 60,      # fast to grow    },)

Two settings deserve attention. MinCapacity=2 rather than 1: a single instance means a single availability zone and no capacity during a rolling update. And the asymmetric cooldowns — scale out in 60 seconds, scale in after 600 — exist because the cost of being slow to add capacity is dropped requests, while the cost of being slow to remove it is a few minutes of an instance you did not need. Symmetric cooldowns produce oscillation: the endpoint scales in during a lull, immediately gets busy, scales out again, and you pay start-up cost repeatedly.

The number 1,200 should come from a load test, not from intuition. If your load test showed a single instance holds p99 under target up to 25 requests per second, that is 1,500 per minute; target 70–80% of it to leave headroom for the two-to-three minutes a new instance takes to become ready.

Auto-scaling does not remove the need for capacity planning; it turns capacity planning into choosing the right target number.

GCP Vertex AI: the same shape, different nouns

Vertex AI separates two concepts that SageMaker blends: a Model is a registered artefact, and an Endpoint is a serving address that can host several models with a traffic split between them. That separation is what makes canary releases a single API call.

Python
from google.cloud import aiplatformaiplatform.init(project="my-project", location="europe-west2")model_v3 = aiplatform.Model.upload(    display_name="defect-classifier-2.6.0",    artifact_uri="gs://my-bucket/models/defectnet/2.6.0/",    serving_container_image_uri=(        "europe-docker.pkg.dev/vertex-ai/prediction/pytorch-cpu.2-3:latest"),)endpoint = aiplatform.Endpoint("projects/.../endpoints/8814...")endpoint.deploy(    model=model_v3,    machine_type="n1-standard-4",    min_replica_count=2, max_replica_count=10,    traffic_percentage=10,            # 10% here, 90% stays on the existing model)

Shifting traffic afterwards touches no infrastructure at all:

Python
# Keys are deployed-model IDs; endpoint.list_models() shows them.endpoint.update(traffic_split={"deployed_model_v25": 50, "deployed_model_v26": 50})# and if the canary metrics look wrong:endpoint.update(traffic_split={"deployed_model_v25": 100, "deployed_model_v26": 0})

That rollback takes a few seconds and does not restart anything, because both models are already loaded and running. Compare with the container-rebuild-and-redeploy path, which is several minutes. During an incident that gap is the whole story.

SageMaker offers the equivalent through production variants with InitialVariantWeight, but the Vertex model where the endpoint owns the split is easier to reason about and harder to get wrong.

Hugging Face Spaces: demos and prototypes

Spaces is a git repository that Hugging Face builds and hosts. Push an app file and a requirements list, and you get a public URL. It is not for production APIs — it sleeps when idle and has no SLA — but it is unmatched for the specific job of putting a working model in front of a non-technical stakeholder within an hour.

Python
# app.py  (Gradio)import gradio as gr, torchimport torchvision.transforms as Tmodel = torch.jit.load("model.pt", map_location="cpu")model.eval()LABELS = ["ok", "scratch", "dent", "discolour"]tfm = T.Compose([T.Resize((224, 224)), T.ToTensor(),                 T.Normalize([0.485, 0.456, 0.406], [0.229, 0.224, 0.225])])def classify(image):    with torch.no_grad():        probs = torch.softmax(model(tfm(image).unsqueeze(0)), dim=1)[0]    return {LABELS[i]: float(probs[i]) for i in range(len(LABELS))}demo = gr.Interface(    fn=classify,    inputs=gr.Image(type="pil", label="Upload a part photo"),    outputs=gr.Label(num_top_classes=4),    examples=[["examples/scratch_01.jpg"], ["examples/ok_04.jpg"]],    title="Surface Defect Classifier",    description="Upload a photo of a machined part.",)if __name__ == "__main__":    demo.launch()
Bash
hf repos create defect-demo --type space --space-sdk gradiogit clone https://huggingface.co/spaces/your-name/defect-demo# add app.py, requirements.txt, model.pt (via git-lfs), examples/git lfs track "*.pt"git add -Agit commit -m "initial demo"git push                      # build starts automatically

Gradio suits single-input-to-single-output demos and generates a REST API automatically. Streamlit suits multi-step apps with sidebars, filters, and state. Both work; pick by the shape of the interaction, not by preference.

Two practical notes. Model files over 10 MB must go through Git LFS (or Xet, the Hub's newer large-file backend, which accepts LFS pushes), and forgetting this is the most common reason a Space fails to build. And the free CPU tier goes to sleep after 48 hours without visitors, so the first visit after a quiet spell has to wait for the Space to restart — fine for a demo, disqualifying for anything a customer touches.

AWS Lambda: pay only for what you use

Lambda runs your function in response to an event and bills per millisecond of execution multiplied by memory allocated. There is no idle cost at all. Container image support allows up to 10 GB, which is enough for most non-LLM models.

Python
import json, base64, io, osimport numpy as np, onnxruntime as ortfrom PIL import Image# Module scope: runs once per cold start, then reused across warm invocations.SESSION = ort.InferenceSession("/opt/ml/model.onnx", providers=["CPUExecutionProvider"])LABELS = ["ok", "scratch", "dent", "discolour"]MEAN = np.array([0.485, 0.456, 0.406], dtype=np.float32)STD  = np.array([0.229, 0.224, 0.225], dtype=np.float32)def handler(event, context):    try:        body = json.loads(event["body"])        raw = base64.b64decode(body["image_b64"])        img = Image.open(io.BytesIO(raw)).convert("RGB").resize((224, 224))        x = ((np.asarray(img, dtype=np.float32) / 255.0 - MEAN) / STD)        x = x.transpose(2, 0, 1)[None, ...]        logits = SESSION.run(None, {"pixels": x})[0][0]        probs = np.exp(logits - logits.max())        probs /= probs.sum()        idx = int(probs.argmax())        return {"statusCode": 200,                "body": json.dumps({"label": LABELS[idx],                                    "confidence": round(float(probs[idx]), 4),                                    "request_id": context.aws_request_id})}    except Exception as e:        return {"statusCode": 500, "body": json.dumps({"error": str(e)})}

The placement of SESSION at module scope is the single most important line. Lambda keeps the execution environment alive between invocations, so module-level code runs on a cold start and is then reused. Move that line inside handler and every invocation reloads the model — turning a 60 ms warm request into a 900 ms one, and multiplying your bill by roughly fifteen.

The cold-start problem, quantified

A cold start for an ML container is not the 100 ms you see in Lambda tutorials:

PhaseTypical duration
Pull and unpack a 3 GB container image2–6 s
Python interpreter and import numpy, PIL, onnxruntime1–3 s
Load the model into memory0.5–3 s
First inference (allocations, kernel selection)0.3–1 s
Total cold start4–13 s
Subsequent warm invocation50–300 ms

Four mitigations, in order of effectiveness. Provisioned concurrency keeps N environments permanently warm — it eliminates cold starts and reintroduces a fixed hourly cost, so it is a partial return to the endpoint model. Shrink the image: ONNX Runtime instead of full PyTorch cuts a 3 GB image to around 400 MB and the pull phase with it. Raise the memory setting: Lambda allocates CPU proportionally to memory, so 3 GB gets roughly twice the CPU of 1.5 GB, and a function that runs twice as fast at twice the memory costs exactly the same while halving latency. A scheduled warm-up ping every five minutes is the cheap hack; it keeps one environment alive but does nothing for concurrent requests.

The cost crossover, worked out

Take a 3 GB Lambda function with an 800 ms warm execution time, at roughly 0.0000166667 USD per GB-second and 0.20 USD per million requests (x86 list prices in us-east-1 at the time of writing; check the current pricing page before you rely on them).

Per invocation: 3 GB × 0.8 s = 2.4 GB-seconds, at 0.0000166667 = 0.00004 USD.

Per million invocations: 40.00 USD of compute plus 0.20 USD of request charges = 40.20 USD.

A SageMaker ml.m5.xlarge endpoint at roughly 0.23 USD per hour (same caveat) runs 24 × 30 = 720 hours = 165.60 USD per month, regardless of traffic.

Break-even: 165.60 / 40.20 = 4.12 million invocations per month, which is 4,120,000 / 2,592,000 seconds = about 1.6 requests per second sustained.

Monthly volumeSustained rateLambdaOne m5.xlarge endpointCheaper
100,0000.04 req/s4.02 USD165.60 USDLambda, by 41×
1,000,0000.4 req/s40.20 USD165.60 USDLambda, by 4×
4,120,0001.6 req/s165.60 USD165.60 USDBreak-even
20,000,0007.7 req/s804 USD165.60 USDEndpoint, by 4.9×
50,000,00019 req/s2,010 USD165.60 USDEndpoint, by 12×

The startup at the top of this lesson sat in the first row and paid the price of a row they were nowhere near. Note also that the break-even moves: an 8 GB, 3-second function breaks even at around 400,000 invocations a month, so the arithmetic must be redone for your actual function, not copied.

Serverless is cheap because you have low traffic, not because it is intrinsically cheap; above a couple of requests per second the economics invert completely.

A decision framework

SituationChooseBecause
Showing a model to stakeholders this weekHugging Face SpacesFree, an hour of work, public URL
Under ~1 req/s, spiky, latency tolerantLambda / Cloud RunNo idle cost; cold starts acceptable
Steady traffic above ~2 req/s, p99 under 200 msSageMaker / Vertex endpointWarm instances, predictable latency
Scoring 40M rows nightlyBatch Transform / Batch PredictionNo endpoint to keep alive; parallel over instances
GPU needed, traffic under 30% duty cycleEndpoint with aggressive auto-scaling, or async inferenceAn idle GPU is the most expensive idle resource there is
Already running KubernetesYour own cluster with KServeNo new platform to learn; portable
Strict data residency or on-premises requirementSelf-hosted containersManaged endpoints constrain where data goes

Two structural considerations sit underneath the table. Managed endpoints impose payload limits — SageMaker real-time caps requests at 25 MB and processing at 60 seconds (8 minutes for streamed responses) — so large images, long videos, or slow generative models need the async or batch variant instead. And every managed platform's model contract is proprietary: your container is portable, but the deployment code, autoscaling configuration, and monitoring integration are not. Keeping inference logic in a plain container with a standard HTTP interface, and treating the platform wrapper as a thin adapter, is what keeps that switching cost from becoming a rewrite.

Mistakes that show up on the bill or the pager

Deploying a GPU endpoint for a model that runs fine on CPU. A quantized ResNet-50 does 60 images per second on a four-core CPU. An ml.g4dn.xlarge endpoint costs roughly three times an ml.m5.xlarge one. Measure CPU latency before you assume you need a GPU; for anything smaller than a large transformer, you usually do not.

Leaving endpoints running. Test and staging endpoints do not stop themselves. Tag every endpoint with an owner and an expiry date, and run a scheduled job that alerts on any endpoint older than its expiry with fewer than a hundred invocations in the past week.

Loading the model inside the handler. Covered above for Lambda; the same error appears in SageMaker when people put loading logic in predict_fn instead of model_fn. The symptom is identical — latency proportional to model size on every single request.

No timeout on the client. A cold Lambda takes eight seconds. If your caller's HTTP timeout is three seconds, it gives up and retries — and the retry hits another cold environment, so you now have two cold starts and one angry user. Set client timeouts above your worst realistic cold start, and use idempotency keys so retries are safe.

Assuming the free tier lasts. A Space that becomes genuinely popular will be rate-limited or asked to move to a paid tier. That is fine when it is a demo. It is an incident when someone has quietly wired a production process to it.

How to make the choice reversible

The decision that actually matters is not which platform you pick first — it is how expensive it is to be wrong. Traffic changes, prices change, and the workload you launch with is rarely the workload you have in a year.

What keeps the choice cheap to revise is where you put the inference logic. Write the model loading, preprocessing, prediction, and postprocessing as an ordinary Python module with no cloud imports in it. Wrap it in a plain FastAPI app in a container. Then the SageMaker inference.py, the Lambda handler, and the Gradio classify function are each ten lines of adapter around the same core, and moving between them is an afternoon rather than a migration project.

Then instrument two numbers from day one and review them monthly: requests per second at peak, and cost per thousand predictions. The first tells you when you have outgrown serverless. The second tells you when you are paying for idle capacity. The startup at the top of this lesson had both numbers available in CloudWatch the entire time — nobody had made looking at them anyone's job.