Course Content
AI System Design and Architecture
3 sections · 7 lessons
Designing Inference APIs — Contracts That Last
A team exposes their sentiment model as POST /predict. The body is {"text": "..."} and the response is a bare float: 0.87. Six weeks and forty integrations later they need to add a second model, so they add POST /predict2. Then a customer asks which model version produced a score, so they change the response to {"score": 0.87, "version": "v3"} — and break every client in production at 09:14 on a Wednesday, because forty clients were all parsing a bare number.
Then a client sends a 900 KB document. The model truncates silently at 512 tokens and returns a confident score for the first paragraph. Nobody notices for a month.
None of this is a modelling failure. It is a contract failure. An inference API is a promise about behaviour that other people build on, and once they have built on it you cannot take it back. Designing it well costs a day; designing it badly costs a year of migration emails.
What makes an inference API different
Most REST guidance assumes CRUD over records. Inference has four properties that break those assumptions:
- The response is probabilistic. The same input can give a different answer next month because the model changed. Clients need to know which model answered, on every single response.
- Latency varies by an order of magnitude with input size. A 20-word request and a 4,000-word request are not the same operation, and a single timeout cannot serve both.
- Cost is proportional to input, not to request count. Rate limiting by requests per minute lets one client consume 80× the resources of another while both stay "within limits".
- Some operations cannot finish inside an HTTP request. Video analysis or document extraction takes minutes. Synchronous is not an option.
Four principles, each with a failure attached
| Principle | Applied | What breaks without it |
|---|---|---|
| Simplicity | One obvious way to do the common thing; required fields minimal | Clients copy a broken example from a forum and you support it forever |
| Consistency | Same envelope, same error shape, same field naming across every endpoint | Every new endpoint costs the client a fresh integration |
| Predictability | Same input, same declared version, same output; documented limits enforced not truncated | Silent truncation — the 900 KB document scored on its first paragraph |
| Extensibility | Objects not scalars; additive changes only; unknown fields ignored | The bare-float outage: no room to add anything |
Never return a bare scalar or a bare array at the top level of a response. An object is the only shape you can add a field to without breaking somebody.
Resources, not verbs
The instinct is to name endpoints after actions: /predict, /classify, /analyse. It works until you have eleven of them and no way to talk about a prediction after it was made. Model the output as a resource instead.
| Verb-shaped (avoid) | Resource-shaped (prefer) | Why |
|---|---|---|
POST /predict | POST /v1/predictions | Creates a prediction; the result has an identity |
POST /predict2 | POST /v1/predictions with model field | New models are data, not new endpoints |
GET /getResult?id=x | GET /v1/predictions/{id} | Cacheable, bookmarkable, standard |
POST /batchPredict | POST /v1/batch-predictions | A batch job is a resource with a lifecycle |
GET /listModels | GET /v1/models | Collections are plural nouns |
POST /v1/predictions create one prediction (sync)GET /v1/predictions/{id} retrieve a stored predictionPOST /v1/batch-predictions create an async batch job → 202GET /v1/batch-predictions/{id} poll job status / resultDELETE /v1/batch-predictions/{id} cancel a running jobGET /v1/models list available models + versionsGET /v1/models/{name} metadata, limits, input schemaGET /healthz /readyz liveness / readiness (unversioned)Request design
Three rules earn their keep. Wrap inputs in a named field rather than putting them at the top level, so options have somewhere to live. Make the model an explicit optional field with a documented default, so a client can pin a version. And accept an idempotency key, so a client retrying after a timeout does not pay twice.
1{2 "input": {"text": "The delivery arrived three days late."},3 "model": "sentiment@v3",4 "options": {"return_probabilities": true, "threshold": 0.5},5 "metadata": {"request_source": "checkout-page"}6}Response design
Every response carries the answer, the identity of what produced it, and enough diagnostics to debug without a support ticket.
1{2 "id": "pred_01HQ8ZK3M9",3 "object": "prediction",4 "created_at": "2026-02-11T09:14:07Z",5 "model": "sentiment@v3",6 "result": {7 "label": "negative",8 "confidence": 0.87,9 "probabilities": {"negative": 0.87, "neutral": 0.10, "positive": 0.03}10 },11 "usage": {"input_tokens": 9, "compute_ms": 21},12 "warnings": []13}The usage block is not decoration. It is what lets a client understand its own bill, what lets you enforce cost-based limits honestly, and what turns "your API is slow" into a conversation about a specific number. The warnings array is where truncation would have surfaced instead of being silent.
Status codes that mean something
| Code | Use for | Should the client retry? |
|---|---|---|
| 200 OK | Synchronous prediction returned | — |
| 202 Accepted | Async job created; body carries job id and poll URL | — |
| 400 Bad Request | Malformed JSON, wrong types | No — fix the request |
| 401 / 403 | Missing credentials / valid credentials without permission | No |
| 404 Not Found | Unknown prediction id or model name | No |
| 413 Payload Too Large | Input exceeds the documented limit | No — split the input |
| 422 Unprocessable | Valid JSON, invalid semantics (empty text, unsupported language) | No |
| 429 Too Many Requests | Rate or quota exceeded; must include Retry-After | Yes, after the stated delay |
| 500 Internal | Unexpected server fault | Yes, with backoff |
| 503 Service Unavailable | Model not loaded, deliberately shedding load | Yes, with backoff |
| 504 Gateway Timeout | Upstream exceeded its budget | Yes, but consider the async endpoint |
The distinction that matters most is 4xx versus 5xx, because well-behaved clients retry 5xx and not 4xx. Returning 500 for a validation failure means every bad request is sent three more times, and a client bug becomes a load event on your GPUs.
Implementation
1from fastapi import FastAPI, Header, HTTPException, Request2from pydantic import BaseModel, Field, field_validator3from typing import Literal, Optional4import time, uuid56app = FastAPI(title="Inference API", version="1.0.0")7MAX_CHARS = 20_00089class PredictionInput(BaseModel):10 text: str = Field(..., min_length=1, max_length=MAX_CHARS)1112class PredictionRequest(BaseModel):13 input: PredictionInput14 model: str = "sentiment@v3"15 options: dict = Field(default_factory=dict)1617 @field_validator("model")18 @classmethod19 def known_model(cls, v):20 if v not in REGISTRY:21 raise ValueError(f"unknown model {v!r}; see GET /v1/models")22 return v2324class Usage(BaseModel):25 input_tokens: int26 compute_ms: int2728class PredictionResponse(BaseModel):29 id: str30 object: Literal["prediction"] = "prediction"31 model: str32 result: dict33 usage: Usage34 warnings: list[str] = []3536@app.post("/v1/predictions", response_model=PredictionResponse, status_code=200)37async def create_prediction(38 req: PredictionRequest,39 idempotency_key: Optional[str] = Header(None, alias="Idempotency-Key"),40):41 if idempotency_key and (cached := IDEMPOTENCY.get(idempotency_key)):42 return cached # same key, same answer, no second charge4344 started = time.perf_counter()45 try:46 result, n_tokens = REGISTRY[req.model].predict(req.input.text, **req.options)47 except ModelNotLoaded:48 raise HTTPException(503, detail="model warming up")4950 body = PredictionResponse(51 id="pred_" + uuid.uuid4().hex[:12],52 model=req.model,53 result=result,54 usage=Usage(input_tokens=n_tokens,55 compute_ms=int((time.perf_counter() - started) * 1000)),56 )57 if idempotency_key:58 IDEMPOTENCY.set(idempotency_key, body, ttl=86_400)59 return bodyTwo details do disproportionate work. max_length=MAX_CHARS converts the silent-truncation bug into a clean 422 with a message the client can act on. The idempotency key converts a timeout — which the client cannot distinguish from a failure — from "possibly charged twice, possibly duplicated" into "safe to retry".
Versioning
| Strategy | Looks like | Strengths | Weaknesses |
|---|---|---|---|
| URL path | /v1/predictions | Visible in logs, trivially routable at the load balancer, easy to curl | Coarse — everything moves at once |
| Header | Accept: application/vnd.api+json; version=2 | Clean URLs, per-resource granularity | Invisible in logs and browsers; easy to forget and get a surprise default |
| Query parameter | /predictions?version=2 | Simplest to add to an existing API | Pollutes cache keys; often stripped by proxies |
Use URL path versioning for the API contract and a separate, independent version for the model. They change on completely different timescales: the wire format may be stable for two years while the model is retrained monthly. Welding them together — /v7/predictions because the model is on its seventh training run — forces a client migration for every retrain.
Draw the line clearly:
| Additive, no version bump | Breaking, requires a new version |
|---|---|
| New optional request field | Removing or renaming a response field |
| New field in the response object | Changing a field's type (string → object) |
| New endpoint | Making an optional field required |
| New enum value in a field documented as open | Changing the meaning of an existing value |
| Better model behind the same declared version alias | Changing the default model when clients pin nothing |
Errors that a client can act on
An error body should tell the caller three things: what went wrong, what to do about it, and how to reference it in a support conversation.
1{2 "error": {3 "type": "invalid_request_error",4 "code": "input_too_long",5 "message": "input.text is 24,318 characters; the maximum is 20,000.",6 "param": "input.text",7 "request_id": "req_01HQ8ZK3M9",8 "docs_url": "https://api.example.com/docs/errors#input_too_long"9 }10}The code field is machine-readable and stable; the message is human-readable and may change wording. Clients branch on code, humans read message. Include request_id on every response — success and failure — and log it alongside every internal span, so a customer email containing one string is enough to reconstruct what happened.
Timeout laddering, with arithmetic
Timeouts must decrease as you go inward, with enough margin for the retries each layer performs. Get this wrong and retries are pure waste.
client 30 s ──────────────────────────────────────► └ edge/CDN 25 s ───────────────────────────────────► └ gateway 20 s ──────────────────────────────► └ API 8 s per attempt, ≤2 attempts + 1 s backoff = 17 s └ model call 6 sCheck the sum: two attempts at 8 s plus 1 s of backoff is 17 s, which fits inside the gateway's 20 s with 3 s of margin. The common mistake is a 12-second per-attempt timeout with two attempts. That is 25 s of possible work under a 20 s gateway budget, so the gateway returns 504 while attempt two is still running — the retry can never succeed, and you have doubled GPU load to produce a guaranteed failure. Every retry must be able to complete inside the caller's remaining budget or it should not be attempted.
Retries also multiply across layers. Three layers each retrying three times turns one client request into up to 27 model invocations during a slowdown, which is how a brownout becomes an outage. Retry at exactly one layer, cap total attempts, and add jitter so retries from thousands of clients do not arrive in lockstep.
Authentication, authorisation, and rate limits
Authentication answers "who is this"; authorisation answers "may they do this". For machine-to-machine inference APIs, an API key in an Authorization header is usually right: simple, revocable, and easy to rotate. Store only a hash of the key — if your database leaks, the keys should be useless. Prefix keys with an identifiable string (sk_live_…) so secret scanners can spot them in a public repository before an attacker does.
Why request-count rate limiting fails for AI
Take a limit of 60 requests per minute. Client A sends 60 requests of 100 tokens: 6,000 tokens a minute. Client B sends 60 requests of 8,000 tokens: 480,000 tokens a minute. Both are "within limits"; B consumes 80× the GPU. Your capacity planning is meaningless and B can starve every other tenant while looking like a model citizen.
Rate limit on the resource you actually spend. A limit that does not correspond to GPU seconds is a limit in name only.
Limit on the resource you actually spend. A token bucket where each request withdraws tokens proportional to its cost handles this cleanly:
1import time23class CostBucket:4 """Token bucket priced in model-tokens, not request counts."""5 def __init__(self, capacity: int, refill_per_sec: float):6 self.capacity = capacity # burst allowance7 self.rate = refill_per_sec # sustained rate8 self.tokens = float(capacity)9 self.updated = time.monotonic()1011 def take(self, cost: int) -> tuple[bool, float]:12 now = time.monotonic()13 self.tokens = min(self.capacity,14 self.tokens + (now - self.updated) * self.rate)15 self.updated = now16 if self.tokens >= cost:17 self.tokens -= cost18 return True, 0.019 deficit = cost - self.tokens20 return False, deficit / self.rate # seconds → Retry-AfterSize it from real capacity rather than a round number. Suppose one GPU replica sustains 12,000 tokens per second and you run 10 replicas at 60% target utilisation: total budget is 10×12,000×0.6=72,000 tokens/s. Split across 40 tenants that is 1,800 tokens/s each. Give each a bucket capacity of 36,000 — twenty seconds of burst — so a client can send a burst of 4 requests at 8,000 tokens without a 429, then settles to the sustained rate. A client that has just spent its whole burst and immediately asks for another 5,000 tokens gets a 429 with Retry-After: 3 (5,000 / 1,800 = 2.8 seconds, rounded up), computed as the actual refill time rather than a guess. A single request larger than the bucket's capacity can never succeed, however long it waits, so reject it up front with a 413 rather than a 429.
Always return the state of the limit on every response, so clients can self-pace instead of discovering the wall:
X-RateLimit-Limit-Tokens: 36000X-RateLimit-Remaining-Tokens: 12480X-RateLimit-Reset: 13Retry-After: 3 (only on 429)Long-running work needs a different shape
Anything that can exceed roughly 30 seconds should not be synchronous. Load balancer idle timeouts, mobile network drops and browser limits all conspire against you, and a dropped connection wastes the compute anyway.
POST /v1/batch-predictions │ 202 Accepted ▼ { "id": "job_7f2", "status": "queued", ┌────────┐ "poll_url": "/v1/batch-predictions/job_7f2" } │ API │──► queue ──► workers ──► result store └────────┘ │ ▼ GET /v1/batch-predictions/job_7f2 → { "status": "running", "progress": {"done": 812, "total": 5000}, "eta_seconds": 96 } → { "status": "succeeded", "result_url": "..." }Offer a webhook as well as polling, because polling at one-second intervals for a ten-minute job costs 600 requests to deliver one result. And return partial progress: a client that can show "812 of 5,000" will wait; a client staring at {"status": "running"} for nine minutes will retry the whole job.
Documentation that stays true
Hand-written API docs drift from reality within weeks. Generate them from the same schema objects the server validates against, so a mismatch is impossible by construction. FastAPI produces an OpenAPI document from the Pydantic models above at /openapi.json, which then generates client libraries in a dozen languages.
What generated docs will not give you, and what you must write by hand: every documented limit with its number (maximum input length, maximum batch size, rate limits per plan), the complete list of error codes with the recommended client action for each, and a genuinely runnable example per endpoint. Then wire the examples into CI so a broken example fails the build.
What this means when you ship one
Write the response envelope before you write the model integration. Every field you might ever want — id, model, usage, warnings — costs nothing to include on day one and is a breaking change to add on day two hundred. The bare-float outage in the opening scenario was decided by a single line of code written in the first week.
Then check three specific things before you expose the endpoint to anyone outside your team. Does every documented limit produce a 4xx rather than silent degradation? Does every response carry a request id and a model version? Do your timeouts and retry counts multiply out to something that fits inside the caller's budget, on paper, with the arithmetic written down?
Two named failure modes to design out from the start. Versioning the API with the model forces every client to migrate whenever you retrain, which trains your clients to pin an old version and never move. Rate limiting by request count gives you a limit that does not correspond to any real resource, so a single heavy tenant can exhaust a fleet while every dashboard says the limits are being respected. Both are cheap to fix on day one and expensive to fix once forty integrations depend on the current behaviour.