LLMOps & Deployment

Course Content

LLMOps & Deployment

6 sections · 40 lessons

What is the gateway pattern for managing LLM APIs?


One door for every model callLLM gatewayVirtual keysand team budgetsAlias to model routingRetries, cooldown, fallbackToken and cost accountingCache and guardrail hooksTraces inOpenTelemetry format
When the looping script hit its team budget, the gateway stopped it — no application had to know.

What you need to know

What it centralises

  • Keys and auth — provider API keys live only in the gateway. Teams get their own virtual keys, which can be revoked.
  • Quotas and budgets — requests per minute, tokens per minute and monthly dollar budgets per team or tenant.
  • Routing — aliases map to one or more deployments; load balance between them.
  • Reliability — retries with backoff, cooldown of failing deployments, fallback to another model.
  • Caching — exact-match response cache; pass-through of provider prompt caching.
  • Accounting — tokens and cost per team, feature and model, for chargeback.
  • Observability and policy — one place to emit traces and run guardrails.

Tools

LiteLLM Proxy (open source) is common; others include Portkey, Kong AI Gateway, Cloudflare AI Gateway, and cloud-native options. A LiteLLM-style config looks like this:

YAML
model_list:  - model_name: chat-default            # the alias apps call    litellm_params:      model: anthropic/claude-sonnet-4-5      api_key: os.environ/ANTHROPIC_API_KEY  - model_name: chat-default            # same alias, second provider    litellm_params:      model: bedrock/anthropic.claude-sonnet-4-5-20250929-v1:0  - model_name: chat-small              # self-hosted vLLM    litellm_params:      model: hosted_vllm/meta-llama/Llama-3.1-8B-Instruct      api_base: http://vllm.internal:8000/v1router_settings:  num_retries: 2  timeout: 30              # seconds per attempt  allowed_fails: 3         # failures before a deployment cools down  cooldown_time: 60  fallbacks: [{"chat-default": ["chat-small"]}]

Two deployments share the chat-default alias, so the gateway spreads load across the direct API and Bedrock. A deployment that fails three times is taken out for 60 seconds. If every chat-default deployment fails, the request falls back to the self-hosted small model.

The trade-offs

  • An extra hop — a few milliseconds, small next to seconds of generation.
  • A new single point of failure — run several replicas, keep it stateless (state in Redis), and health-check it.
  • Buffering bugs — the gateway must pass streamed tokens through immediately.

A real-life example

A company's internal code assistant is used by 2,000 engineers across 40 teams, and other teams start building their own AI tools. Before the gateway, six services each held a provider key, and a leaked key in a CI log cost a weekend of rotation.

After the gateway: each team gets a virtual key with a monthly budget (say $1,500 for most teams, more for the assistant itself). When one team's nightly script loops and burns $900 in an hour, the gateway's budget stops it at the limit and alerts the owner. When the provider has an outage, the fallback to the self-hosted model keeps the assistant answering simple questions. And when the platform team moves to a newer model, they change one alias in one file.

Follow-up questions to expect

  • "Where do you enforce rate limits — gateway or app?" — At the gateway, keyed by team or tenant, with shared state in Redis so limits hold across replicas.
  • "Does the gateway break prompt caching?" — No, as long as it forwards the prompt unchanged; provider caches key on the prompt prefix, not on the client.
  • "Build or buy?" — Start with an open-source proxy; build only the pieces that are specific to you, such as custom guardrails.