AI System Design and Architecture

Designing Inference APIs — Contracts That Last


A team exposes their sentiment model as POST /predict. The body is {"text": "..."} and the response is a bare float: 0.87. Six weeks and forty integrations later they need to add a second model, so they add POST /predict2. Then a customer asks which model version produced a score, so they change the response to {"score": 0.87, "version": "v3"} — and break every client in production at 09:14 on a Wednesday, because forty clients were all parsing a bare number.

Then a client sends a 900 KB document. The model truncates silently at 512 tokens and returns a confident score for the first paragraph. Nobody notices for a month.

None of this is a modelling failure. It is a contract failure. An inference API is a promise about behaviour that other people build on, and once they have built on it you cannot take it back. Designing it well costs a day; designing it badly costs a year of migration emails.

A bare float has nowhere for the next field to goThe contract that rotted• POST /predict returns0.87 and nothing else• A second model becomes POST /predict2• Adding a version field breaks 40 clients• No way to say which model scored itA contract that survives• POST /v1/sentiment/predictions• An object with named, documented fields• model_version present in every reply• Only additive changes inside a version
Every later requirement — versions, batching, confidence, explanations — needs a field, and a scalar has none.

What makes an inference API different

Most REST guidance assumes CRUD over records. Inference has four properties that break those assumptions:

  • The response is probabilistic. The same input can give a different answer next month because the model changed. Clients need to know which model answered, on every single response.
  • Latency varies by an order of magnitude with input size. A 20-word request and a 4,000-word request are not the same operation, and a single timeout cannot serve both.
  • Cost is proportional to input, not to request count. Rate limiting by requests per minute lets one client consume 80× the resources of another while both stay "within limits".
  • Some operations cannot finish inside an HTTP request. Video analysis or document extraction takes minutes. Synchronous is not an option.

Four principles, each with a failure attached

PrincipleAppliedWhat breaks without it
SimplicityOne obvious way to do the common thing; required fields minimalClients copy a broken example from a forum and you support it forever
ConsistencySame envelope, same error shape, same field naming across every endpointEvery new endpoint costs the client a fresh integration
PredictabilitySame input, same declared version, same output; documented limits enforced not truncatedSilent truncation — the 900 KB document scored on its first paragraph
ExtensibilityObjects not scalars; additive changes only; unknown fields ignoredThe bare-float outage: no room to add anything

Never return a bare scalar or a bare array at the top level of a response. An object is the only shape you can add a field to without breaking somebody.

Resources, not verbs

The instinct is to name endpoints after actions: /predict, /classify, /analyse. It works until you have eleven of them and no way to talk about a prediction after it was made. Model the output as a resource instead.

Verb-shaped (avoid)Resource-shaped (prefer)Why
POST /predictPOST /v1/predictionsCreates a prediction; the result has an identity
POST /predict2POST /v1/predictions with model fieldNew models are data, not new endpoints
GET /getResult?id=xGET /v1/predictions/{id}Cacheable, bookmarkable, standard
POST /batchPredictPOST /v1/batch-predictionsA batch job is a resource with a lifecycle
GET /listModelsGET /v1/modelsCollections are plural nouns
Text
POST   /v1/predictions              create one prediction (sync)GET    /v1/predictions/{id}         retrieve a stored predictionPOST   /v1/batch-predictions        create an async batch job → 202GET    /v1/batch-predictions/{id}   poll job status / resultDELETE /v1/batch-predictions/{id}   cancel a running jobGET    /v1/models                   list available models + versionsGET    /v1/models/{name}            metadata, limits, input schemaGET    /healthz   /readyz           liveness / readiness (unversioned)

Request design

Three rules earn their keep. Wrap inputs in a named field rather than putting them at the top level, so options have somewhere to live. Make the model an explicit optional field with a documented default, so a client can pin a version. And accept an idempotency key, so a client retrying after a timeout does not pay twice.

JSON
{  "input": {"text": "The delivery arrived three days late."},  "model": "sentiment@v3",  "options": {"return_probabilities": true, "threshold": 0.5},  "metadata": {"request_source": "checkout-page"}}

Response design

Every response carries the answer, the identity of what produced it, and enough diagnostics to debug without a support ticket.

JSON
{  "id": "pred_01HQ8ZK3M9",  "object": "prediction",  "created_at": "2026-02-11T09:14:07Z",  "model": "sentiment@v3",  "result": {    "label": "negative",    "confidence": 0.87,    "probabilities": {"negative": 0.87, "neutral": 0.10, "positive": 0.03}  },  "usage": {"input_tokens": 9, "compute_ms": 21},  "warnings": []}

The usage block is not decoration. It is what lets a client understand its own bill, what lets you enforce cost-based limits honestly, and what turns "your API is slow" into a conversation about a specific number. The warnings array is where truncation would have surfaced instead of being silent.

Status codes that mean something

CodeUse forShould the client retry?
200 OKSynchronous prediction returned—
202 AcceptedAsync job created; body carries job id and poll URL—
400 Bad RequestMalformed JSON, wrong typesNo — fix the request
401 / 403Missing credentials / valid credentials without permissionNo
404 Not FoundUnknown prediction id or model nameNo
413 Payload Too LargeInput exceeds the documented limitNo — split the input
422 UnprocessableValid JSON, invalid semantics (empty text, unsupported language)No
429 Too Many RequestsRate or quota exceeded; must include Retry-AfterYes, after the stated delay
500 InternalUnexpected server faultYes, with backoff
503 Service UnavailableModel not loaded, deliberately shedding loadYes, with backoff
504 Gateway TimeoutUpstream exceeded its budgetYes, but consider the async endpoint

The distinction that matters most is 4xx versus 5xx, because well-behaved clients retry 5xx and not 4xx. Returning 500 for a validation failure means every bad request is sent three more times, and a client bug becomes a load event on your GPUs.

Implementation

Python
from fastapi import FastAPI, Header, HTTPException, Requestfrom pydantic import BaseModel, Field, field_validatorfrom typing import Literal, Optionalimport time, uuidapp = FastAPI(title="Inference API", version="1.0.0")MAX_CHARS = 20_000class PredictionInput(BaseModel):    text: str = Field(..., min_length=1, max_length=MAX_CHARS)class PredictionRequest(BaseModel):    input: PredictionInput    model: str = "sentiment@v3"    options: dict = Field(default_factory=dict)    @field_validator("model")    @classmethod    def known_model(cls, v):        if v not in REGISTRY:            raise ValueError(f"unknown model {v!r}; see GET /v1/models")        return vclass Usage(BaseModel):    input_tokens: int    compute_ms: intclass PredictionResponse(BaseModel):    id: str    object: Literal["prediction"] = "prediction"    model: str    result: dict    usage: Usage    warnings: list[str] = []@app.post("/v1/predictions", response_model=PredictionResponse, status_code=200)async def create_prediction(    req: PredictionRequest,    idempotency_key: Optional[str] = Header(None, alias="Idempotency-Key"),):    if idempotency_key and (cached := IDEMPOTENCY.get(idempotency_key)):        return cached                      # same key, same answer, no second charge    started = time.perf_counter()    try:        result, n_tokens = REGISTRY[req.model].predict(req.input.text, **req.options)    except ModelNotLoaded:        raise HTTPException(503, detail="model warming up")    body = PredictionResponse(        id="pred_" + uuid.uuid4().hex[:12],        model=req.model,        result=result,        usage=Usage(input_tokens=n_tokens,                    compute_ms=int((time.perf_counter() - started) * 1000)),    )    if idempotency_key:        IDEMPOTENCY.set(idempotency_key, body, ttl=86_400)    return body

Two details do disproportionate work. max_length=MAX_CHARS converts the silent-truncation bug into a clean 422 with a message the client can act on. The idempotency key converts a timeout — which the client cannot distinguish from a failure — from "possibly charged twice, possibly duplicated" into "safe to retry".

Versioning

StrategyLooks likeStrengthsWeaknesses
URL path/v1/predictionsVisible in logs, trivially routable at the load balancer, easy to curlCoarse — everything moves at once
HeaderAccept: application/vnd.api+json; version=2Clean URLs, per-resource granularityInvisible in logs and browsers; easy to forget and get a surprise default
Query parameter/predictions?version=2Simplest to add to an existing APIPollutes cache keys; often stripped by proxies

Use URL path versioning for the API contract and a separate, independent version for the model. They change on completely different timescales: the wire format may be stable for two years while the model is retrained monthly. Welding them together — /v7/predictions because the model is on its seventh training run — forces a client migration for every retrain.

Draw the line clearly:

Additive, no version bumpBreaking, requires a new version
New optional request fieldRemoving or renaming a response field
New field in the response objectChanging a field's type (string → object)
New endpointMaking an optional field required
New enum value in a field documented as openChanging the meaning of an existing value
Better model behind the same declared version aliasChanging the default model when clients pin nothing

Errors that a client can act on

An error body should tell the caller three things: what went wrong, what to do about it, and how to reference it in a support conversation.

JSON
{  "error": {    "type": "invalid_request_error",    "code": "input_too_long",    "message": "input.text is 24,318 characters; the maximum is 20,000.",    "param": "input.text",    "request_id": "req_01HQ8ZK3M9",    "docs_url": "https://api.example.com/docs/errors#input_too_long"  }}

The code field is machine-readable and stable; the message is human-readable and may change wording. Clients branch on code, humans read message. Include request_id on every response — success and failure — and log it alongside every internal span, so a customer email containing one string is enough to reconstruct what happened.

Timeout laddering, with arithmetic

Timeouts must decrease as you go inward, with enough margin for the retries each layer performs. Get this wrong and retries are pure waste.

Text
  client        30 s  ──────────────────────────────────────►   └ edge/CDN   25 s  ───────────────────────────────────►      └ gateway 20 s  ──────────────────────────────►         └ API   8 s per attempt, ≤2 attempts + 1 s backoff = 17 s            └ model call 6 s

Check the sum: two attempts at 8 s plus 1 s of backoff is 17 s, which fits inside the gateway's 20 s with 3 s of margin. The common mistake is a 12-second per-attempt timeout with two attempts. That is 25 s of possible work under a 20 s gateway budget, so the gateway returns 504 while attempt two is still running — the retry can never succeed, and you have doubled GPU load to produce a guaranteed failure. Every retry must be able to complete inside the caller's remaining budget or it should not be attempted.

Retries also multiply across layers. Three layers each retrying three times turns one client request into up to 27 model invocations during a slowdown, which is how a brownout becomes an outage. Retry at exactly one layer, cap total attempts, and add jitter so retries from thousands of clients do not arrive in lockstep.

Authentication, authorisation, and rate limits

Authentication answers "who is this"; authorisation answers "may they do this". For machine-to-machine inference APIs, an API key in an Authorization header is usually right: simple, revocable, and easy to rotate. Store only a hash of the key — if your database leaks, the keys should be useless. Prefix keys with an identifiable string (sk_live_…) so secret scanners can spot them in a public repository before an attacker does.

Why request-count rate limiting fails for AI

Take a limit of 60 requests per minute. Client A sends 60 requests of 100 tokens: 6,000 tokens a minute. Client B sends 60 requests of 8,000 tokens: 480,000 tokens a minute. Both are "within limits"; B consumes 80× the GPU. Your capacity planning is meaningless and B can starve every other tenant while looking like a model citizen.

Rate limit on the resource you actually spend. A limit that does not correspond to GPU seconds is a limit in name only.

Limit on the resource you actually spend. A token bucket where each request withdraws tokens proportional to its cost handles this cleanly:

Python
import timeclass CostBucket:    """Token bucket priced in model-tokens, not request counts."""    def __init__(self, capacity: int, refill_per_sec: float):        self.capacity = capacity          # burst allowance        self.rate = refill_per_sec        # sustained rate        self.tokens = float(capacity)        self.updated = time.monotonic()    def take(self, cost: int) -> tuple[bool, float]:        now = time.monotonic()        self.tokens = min(self.capacity,                          self.tokens + (now - self.updated) * self.rate)        self.updated = now        if self.tokens >= cost:            self.tokens -= cost            return True, 0.0        deficit = cost - self.tokens        return False, deficit / self.rate     # seconds → Retry-After

Size it from real capacity rather than a round number. Suppose one GPU replica sustains 12,000 tokens per second and you run 10 replicas at 60% target utilisation: total budget is 10×12,000×0.6=72,00010 \times 12{,}000 \times 0.6 = 72{,}000 tokens/s. Split across 40 tenants that is 1,800 tokens/s each. Give each a bucket capacity of 36,000 — twenty seconds of burst — so a client can send a burst of 4 requests at 8,000 tokens without a 429, then settles to the sustained rate. A client that has just spent its whole burst and immediately asks for another 5,000 tokens gets a 429 with Retry-After: 3 (5,000 / 1,800 = 2.8 seconds, rounded up), computed as the actual refill time rather than a guess. A single request larger than the bucket's capacity can never succeed, however long it waits, so reject it up front with a 413 rather than a 429.

Always return the state of the limit on every response, so clients can self-pace instead of discovering the wall:

Text
X-RateLimit-Limit-Tokens:     36000X-RateLimit-Remaining-Tokens: 12480X-RateLimit-Reset:            13Retry-After:                  3          (only on 429)

Long-running work needs a different shape

Anything that can exceed roughly 30 seconds should not be synchronous. Load balancer idle timeouts, mobile network drops and browser limits all conspire against you, and a dropped connection wastes the compute anyway.

Text
  POST /v1/batch-predictions       │                     202 Accepted       ▼                     { "id": "job_7f2", "status": "queued",   ┌────────┐                  "poll_url": "/v1/batch-predictions/job_7f2" }   │  API   │──► queue ──► workers ──► result store   └────────┘                                │                                             ▼  GET /v1/batch-predictions/job_7f2   →  { "status": "running",                                           "progress": {"done": 812,                                                        "total": 5000},                                           "eta_seconds": 96 }                                      →  { "status": "succeeded",                                           "result_url": "..." }

Offer a webhook as well as polling, because polling at one-second intervals for a ten-minute job costs 600 requests to deliver one result. And return partial progress: a client that can show "812 of 5,000" will wait; a client staring at {"status": "running"} for nine minutes will retry the whole job.

Documentation that stays true

Hand-written API docs drift from reality within weeks. Generate them from the same schema objects the server validates against, so a mismatch is impossible by construction. FastAPI produces an OpenAPI document from the Pydantic models above at /openapi.json, which then generates client libraries in a dozen languages.

What generated docs will not give you, and what you must write by hand: every documented limit with its number (maximum input length, maximum batch size, rate limits per plan), the complete list of error codes with the recommended client action for each, and a genuinely runnable example per endpoint. Then wire the examples into CI so a broken example fails the build.

What this means when you ship one

Write the response envelope before you write the model integration. Every field you might ever want — id, model, usage, warnings — costs nothing to include on day one and is a breaking change to add on day two hundred. The bare-float outage in the opening scenario was decided by a single line of code written in the first week.

Then check three specific things before you expose the endpoint to anyone outside your team. Does every documented limit produce a 4xx rather than silent degradation? Does every response carry a request id and a model version? Do your timeouts and retry counts multiply out to something that fits inside the caller's budget, on paper, with the arithmetic written down?

Two named failure modes to design out from the start. Versioning the API with the model forces every client to migrate whenever you retrain, which trains your clients to pin an old version and never move. Rate limiting by request count gives you a limit that does not correspond to any real resource, so a single heavy tenant can exhaust a fleet while every dashboard says the limits are being respected. Both are cheap to fix on day one and expensive to fix once forty integrations depend on the current behaviour.