Course Content
Model Deployment for AI Engineers
4 sections · 10 lessons
Cloud Deployment (AWS SageMaker, GCP Vertex AI, Hugging Face Spaces, AWS Lambda)
A startup deploys an image classifier to a SageMaker real-time endpoint on an ml.g4dn.xlarge GPU instance. It works beautifully. Six weeks later finance flags a cloud bill of roughly 540 USD for that one endpoint. The traffic logs show 8,400 predictions that month — about twelve per hour. They paid for 720 hours of an idle GPU to serve twelve requests an hour.
The same workload on AWS Lambda would have cost under 1 USD. The model was fine, the code was fine, the platform was fine — the match between traffic shape and pricing model was catastrophically wrong.
That is the recurring theme of cloud deployment. These platforms are not ranked best to worst. They are priced and engineered for different traffic patterns, and picking the wrong one costs you either an order of magnitude in money or an order of magnitude in latency. This lesson is about telling them apart.
Four different things called "deployment"
| Pattern | You pay for | Latency | Fits |
|---|---|---|---|
| Managed endpoint SageMaker, Vertex AI | Instance-hours, whether or not requests arrive | Consistently 50–200 ms | Steady traffic above roughly 2 req/s |
| Serverless function Lambda, Cloud Run | Invocations and GB-seconds only | 50 ms warm, 3–15 s cold | Spiky, low, or unpredictable traffic |
| Batch / async job Batch Transform, Batch Prediction | Compute for the duration of the job | Minutes to hours | Scoring millions of rows overnight |
| Hosted demo Hugging Face Spaces | Nothing, on the free tier | Seconds, sleeps when idle | Prototypes, stakeholder review, public demos |
Before choosing anything, answer three questions with actual numbers: how many requests per second at peak, how many per month in total, and what latency does the user experience require? Those three numbers determine the answer almost entirely. Teams that skip them tend to reach for whichever platform they read about most recently.
AWS SageMaker: managed endpoints
The mental model: you hand SageMaker a container image and a model artefact in S3. It runs the container on instances it manages, puts a load balancer and TLS in front, and gives you an HTTPS endpoint. The container must implement two routes — GET /ping for health and POST /invocations for inference — and listen on port 8080. That is the entire contract, which means any FastAPI server you already have becomes a SageMaker container with almost no change.
The deployment code below uses the SageMaker Python SDK v2 (pip install "sagemaker<3"). SDK v3 replaced the PyTorchModel-style classes with a single ModelBuilder and does not run v2 code unchanged, so check which major version your project pins. The container contract and the four inference hooks are the same either way.
1import boto3, sagemaker2from sagemaker.pytorch import PyTorchModel34role = "arn:aws:iam::123456789012:role/SageMakerExecutionRole"56model = PyTorchModel(7 model_data="s3://my-bucket/models/defectnet/2.5.0/model.tar.gz",8 role=role,9 entry_point="inference.py", # defines the four handler functions below10 source_dir="code/",11 framework_version="2.3",12 py_version="py311",13 env={"OMP_NUM_THREADS": "1", "MODEL_VERSION": "2.5.0"},14)1516predictor = model.deploy(17 initial_instance_count=2, # two for availability, not for throughput18 instance_type="ml.m5.xlarge",19 endpoint_name="defect-classifier-prod",20)The artefact must be a .tar.gz, and SageMaker extracts it into /opt/ml/model inside the container. Your inference.py defines four hooks:
1import torch, json, io2from PIL import Image3import torchvision.transforms as T45def model_fn(model_dir): # called ONCE at container start6 m = torch.jit.load(f"{model_dir}/model.pt", map_location="cpu")7 m.eval()8 return m910def input_fn(request_body, content_type): # deserialise11 if content_type == "application/x-image":12 img = Image.open(io.BytesIO(request_body)).convert("RGB")13 tfm = T.Compose([T.Resize((224, 224)), T.ToTensor(),14 T.Normalize([0.485, 0.456, 0.406], [0.229, 0.224, 0.225])])15 return tfm(img).unsqueeze(0)16 raise ValueError(f"unsupported content type: {content_type}")1718def predict_fn(x, model): # inference19 with torch.no_grad():20 return torch.softmax(model(x), dim=1)2122def output_fn(probs, accept): # serialise23 idx = int(probs.argmax())24 return json.dumps({"label": LABELS[idx],25 "confidence": round(float(probs[0, idx]), 4)}), "application/json"Splitting the pipeline into these four functions is not arbitrary. It lets SageMaker call model_fn exactly once per container while calling the other three per request — which is the same "load at startup, not in the handler" rule that governs every serving stack, enforced by the framework.
Auto-scaling
A fixed instance count is either wasteful at night or insufficient at peak. Target tracking adjusts the count to hold a chosen metric at a chosen value.
1aas = boto3.client("application-autoscaling")2resource_id = "endpoint/defect-classifier-prod/variant/AllTraffic"34aas.register_scalable_target(5 ServiceNamespace="sagemaker", ResourceId=resource_id,6 ScalableDimension="sagemaker:variant:DesiredInstanceCount",7 MinCapacity=2, MaxCapacity=12,8)910aas.put_scaling_policy(11 PolicyName="invocations-per-instance",12 ServiceNamespace="sagemaker", ResourceId=resource_id,13 ScalableDimension="sagemaker:variant:DesiredInstanceCount",14 PolicyType="TargetTrackingScaling",15 TargetTrackingScalingPolicyConfiguration={16 "TargetValue": 1200.0, # invocations per instance per minute17 "PredefinedMetricSpecification": {18 "PredefinedMetricType": "SageMakerVariantInvocationsPerInstance"19 },20 "ScaleInCooldown": 600, # slow to shrink21 "ScaleOutCooldown": 60, # fast to grow22 },23)Two settings deserve attention. MinCapacity=2 rather than 1: a single instance means a single availability zone and no capacity during a rolling update. And the asymmetric cooldowns — scale out in 60 seconds, scale in after 600 — exist because the cost of being slow to add capacity is dropped requests, while the cost of being slow to remove it is a few minutes of an instance you did not need. Symmetric cooldowns produce oscillation: the endpoint scales in during a lull, immediately gets busy, scales out again, and you pay start-up cost repeatedly.
The number 1,200 should come from a load test, not from intuition. If your load test showed a single instance holds p99 under target up to 25 requests per second, that is 1,500 per minute; target 70–80% of it to leave headroom for the two-to-three minutes a new instance takes to become ready.
Auto-scaling does not remove the need for capacity planning; it turns capacity planning into choosing the right target number.
GCP Vertex AI: the same shape, different nouns
Vertex AI separates two concepts that SageMaker blends: a Model is a registered artefact, and an Endpoint is a serving address that can host several models with a traffic split between them. That separation is what makes canary releases a single API call.
1from google.cloud import aiplatform23aiplatform.init(project="my-project", location="europe-west2")45model_v3 = aiplatform.Model.upload(6 display_name="defect-classifier-2.6.0",7 artifact_uri="gs://my-bucket/models/defectnet/2.6.0/",8 serving_container_image_uri=(9 "europe-docker.pkg.dev/vertex-ai/prediction/pytorch-cpu.2-3:latest"),10)1112endpoint = aiplatform.Endpoint("projects/.../endpoints/8814...")1314endpoint.deploy(15 model=model_v3,16 machine_type="n1-standard-4",17 min_replica_count=2, max_replica_count=10,18 traffic_percentage=10, # 10% here, 90% stays on the existing model19)Shifting traffic afterwards touches no infrastructure at all:
1# Keys are deployed-model IDs; endpoint.list_models() shows them.2endpoint.update(traffic_split={"deployed_model_v25": 50, "deployed_model_v26": 50})3# and if the canary metrics look wrong:4endpoint.update(traffic_split={"deployed_model_v25": 100, "deployed_model_v26": 0})That rollback takes a few seconds and does not restart anything, because both models are already loaded and running. Compare with the container-rebuild-and-redeploy path, which is several minutes. During an incident that gap is the whole story.
SageMaker offers the equivalent through production variants with InitialVariantWeight, but the Vertex model where the endpoint owns the split is easier to reason about and harder to get wrong.
Hugging Face Spaces: demos and prototypes
Spaces is a git repository that Hugging Face builds and hosts. Push an app file and a requirements list, and you get a public URL. It is not for production APIs — it sleeps when idle and has no SLA — but it is unmatched for the specific job of putting a working model in front of a non-technical stakeholder within an hour.
1# app.py (Gradio)2import gradio as gr, torch3import torchvision.transforms as T45model = torch.jit.load("model.pt", map_location="cpu")6model.eval()7LABELS = ["ok", "scratch", "dent", "discolour"]8tfm = T.Compose([T.Resize((224, 224)), T.ToTensor(),9 T.Normalize([0.485, 0.456, 0.406], [0.229, 0.224, 0.225])])1011def classify(image):12 with torch.no_grad():13 probs = torch.softmax(model(tfm(image).unsqueeze(0)), dim=1)[0]14 return {LABELS[i]: float(probs[i]) for i in range(len(LABELS))}1516demo = gr.Interface(17 fn=classify,18 inputs=gr.Image(type="pil", label="Upload a part photo"),19 outputs=gr.Label(num_top_classes=4),20 examples=[["examples/scratch_01.jpg"], ["examples/ok_04.jpg"]],21 title="Surface Defect Classifier",22 description="Upload a photo of a machined part.",23)2425if __name__ == "__main__":26 demo.launch()1hf repos create defect-demo --type space --space-sdk gradio2git clone https://huggingface.co/spaces/your-name/defect-demo3# add app.py, requirements.txt, model.pt (via git-lfs), examples/4git lfs track "*.pt"5git add -A6git commit -m "initial demo"7git push # build starts automaticallyGradio suits single-input-to-single-output demos and generates a REST API automatically. Streamlit suits multi-step apps with sidebars, filters, and state. Both work; pick by the shape of the interaction, not by preference.
Two practical notes. Model files over 10 MB must go through Git LFS (or Xet, the Hub's newer large-file backend, which accepts LFS pushes), and forgetting this is the most common reason a Space fails to build. And the free CPU tier goes to sleep after 48 hours without visitors, so the first visit after a quiet spell has to wait for the Space to restart — fine for a demo, disqualifying for anything a customer touches.
AWS Lambda: pay only for what you use
Lambda runs your function in response to an event and bills per millisecond of execution multiplied by memory allocated. There is no idle cost at all. Container image support allows up to 10 GB, which is enough for most non-LLM models.
1import json, base64, io, os2import numpy as np, onnxruntime as ort3from PIL import Image45# Module scope: runs once per cold start, then reused across warm invocations.6SESSION = ort.InferenceSession("/opt/ml/model.onnx", providers=["CPUExecutionProvider"])7LABELS = ["ok", "scratch", "dent", "discolour"]8MEAN = np.array([0.485, 0.456, 0.406], dtype=np.float32)9STD = np.array([0.229, 0.224, 0.225], dtype=np.float32)1011def handler(event, context):12 try:13 body = json.loads(event["body"])14 raw = base64.b64decode(body["image_b64"])15 img = Image.open(io.BytesIO(raw)).convert("RGB").resize((224, 224))16 x = ((np.asarray(img, dtype=np.float32) / 255.0 - MEAN) / STD)17 x = x.transpose(2, 0, 1)[None, ...]1819 logits = SESSION.run(None, {"pixels": x})[0][0]20 probs = np.exp(logits - logits.max())21 probs /= probs.sum()22 idx = int(probs.argmax())2324 return {"statusCode": 200,25 "body": json.dumps({"label": LABELS[idx],26 "confidence": round(float(probs[idx]), 4),27 "request_id": context.aws_request_id})}28 except Exception as e:29 return {"statusCode": 500, "body": json.dumps({"error": str(e)})}The placement of SESSION at module scope is the single most important line. Lambda keeps the execution environment alive between invocations, so module-level code runs on a cold start and is then reused. Move that line inside handler and every invocation reloads the model — turning a 60 ms warm request into a 900 ms one, and multiplying your bill by roughly fifteen.
The cold-start problem, quantified
A cold start for an ML container is not the 100 ms you see in Lambda tutorials:
| Phase | Typical duration |
|---|---|
| Pull and unpack a 3 GB container image | 2–6 s |
| Python interpreter and import numpy, PIL, onnxruntime | 1–3 s |
| Load the model into memory | 0.5–3 s |
| First inference (allocations, kernel selection) | 0.3–1 s |
| Total cold start | 4–13 s |
| Subsequent warm invocation | 50–300 ms |
Four mitigations, in order of effectiveness. Provisioned concurrency keeps N environments permanently warm — it eliminates cold starts and reintroduces a fixed hourly cost, so it is a partial return to the endpoint model. Shrink the image: ONNX Runtime instead of full PyTorch cuts a 3 GB image to around 400 MB and the pull phase with it. Raise the memory setting: Lambda allocates CPU proportionally to memory, so 3 GB gets roughly twice the CPU of 1.5 GB, and a function that runs twice as fast at twice the memory costs exactly the same while halving latency. A scheduled warm-up ping every five minutes is the cheap hack; it keeps one environment alive but does nothing for concurrent requests.
The cost crossover, worked out
Take a 3 GB Lambda function with an 800 ms warm execution time, at roughly 0.0000166667 USD per GB-second and 0.20 USD per million requests (x86 list prices in us-east-1 at the time of writing; check the current pricing page before you rely on them).
Per invocation: 3 GB × 0.8 s = 2.4 GB-seconds, at 0.0000166667 = 0.00004 USD.
Per million invocations: 40.00 USD of compute plus 0.20 USD of request charges = 40.20 USD.
A SageMaker ml.m5.xlarge endpoint at roughly 0.23 USD per hour (same caveat) runs 24 × 30 = 720 hours = 165.60 USD per month, regardless of traffic.
Break-even: 165.60 / 40.20 = 4.12 million invocations per month, which is 4,120,000 / 2,592,000 seconds = about 1.6 requests per second sustained.
| Monthly volume | Sustained rate | Lambda | One m5.xlarge endpoint | Cheaper |
|---|---|---|---|---|
| 100,000 | 0.04 req/s | 4.02 USD | 165.60 USD | Lambda, by 41× |
| 1,000,000 | 0.4 req/s | 40.20 USD | 165.60 USD | Lambda, by 4× |
| 4,120,000 | 1.6 req/s | 165.60 USD | 165.60 USD | Break-even |
| 20,000,000 | 7.7 req/s | 804 USD | 165.60 USD | Endpoint, by 4.9× |
| 50,000,000 | 19 req/s | 2,010 USD | 165.60 USD | Endpoint, by 12× |
The startup at the top of this lesson sat in the first row and paid the price of a row they were nowhere near. Note also that the break-even moves: an 8 GB, 3-second function breaks even at around 400,000 invocations a month, so the arithmetic must be redone for your actual function, not copied.
Serverless is cheap because you have low traffic, not because it is intrinsically cheap; above a couple of requests per second the economics invert completely.
A decision framework
| Situation | Choose | Because |
|---|---|---|
| Showing a model to stakeholders this week | Hugging Face Spaces | Free, an hour of work, public URL |
| Under ~1 req/s, spiky, latency tolerant | Lambda / Cloud Run | No idle cost; cold starts acceptable |
| Steady traffic above ~2 req/s, p99 under 200 ms | SageMaker / Vertex endpoint | Warm instances, predictable latency |
| Scoring 40M rows nightly | Batch Transform / Batch Prediction | No endpoint to keep alive; parallel over instances |
| GPU needed, traffic under 30% duty cycle | Endpoint with aggressive auto-scaling, or async inference | An idle GPU is the most expensive idle resource there is |
| Already running Kubernetes | Your own cluster with KServe | No new platform to learn; portable |
| Strict data residency or on-premises requirement | Self-hosted containers | Managed endpoints constrain where data goes |
Two structural considerations sit underneath the table. Managed endpoints impose payload limits — SageMaker real-time caps requests at 25 MB and processing at 60 seconds (8 minutes for streamed responses) — so large images, long videos, or slow generative models need the async or batch variant instead. And every managed platform's model contract is proprietary: your container is portable, but the deployment code, autoscaling configuration, and monitoring integration are not. Keeping inference logic in a plain container with a standard HTTP interface, and treating the platform wrapper as a thin adapter, is what keeps that switching cost from becoming a rewrite.
Mistakes that show up on the bill or the pager
Deploying a GPU endpoint for a model that runs fine on CPU. A quantized ResNet-50 does 60 images per second on a four-core CPU. An ml.g4dn.xlarge endpoint costs roughly three times an ml.m5.xlarge one. Measure CPU latency before you assume you need a GPU; for anything smaller than a large transformer, you usually do not.
Leaving endpoints running. Test and staging endpoints do not stop themselves. Tag every endpoint with an owner and an expiry date, and run a scheduled job that alerts on any endpoint older than its expiry with fewer than a hundred invocations in the past week.
Loading the model inside the handler. Covered above for Lambda; the same error appears in SageMaker when people put loading logic in predict_fn instead of model_fn. The symptom is identical — latency proportional to model size on every single request.
No timeout on the client. A cold Lambda takes eight seconds. If your caller's HTTP timeout is three seconds, it gives up and retries — and the retry hits another cold environment, so you now have two cold starts and one angry user. Set client timeouts above your worst realistic cold start, and use idempotency keys so retries are safe.
Assuming the free tier lasts. A Space that becomes genuinely popular will be rate-limited or asked to move to a paid tier. That is fine when it is a demo. It is an incident when someone has quietly wired a production process to it.
How to make the choice reversible
The decision that actually matters is not which platform you pick first — it is how expensive it is to be wrong. Traffic changes, prices change, and the workload you launch with is rarely the workload you have in a year.
What keeps the choice cheap to revise is where you put the inference logic. Write the model loading, preprocessing, prediction, and postprocessing as an ordinary Python module with no cloud imports in it. Wrap it in a plain FastAPI app in a container. Then the SageMaker inference.py, the Lambda handler, and the Gradio classify function are each ten lines of adapter around the same core, and moving between them is an afternoon rather than a migration project.
Then instrument two numbers from day one and review them monthly: requests per second at peak, and cost per thousand predictions. The first tells you when you have outgrown serverless. The second tells you when you are paying for idle capacity. The startup at the top of this lesson had both numbers available in CloudWatch the entire time — nobody had made looking at them anyone's job.