Course Content
LLMOps & Deployment
6 sections · 40 lessons
What is the gateway pattern for managing LLM APIs?
What you need to know
What it centralises
- Keys and auth — provider API keys live only in the gateway. Teams get their own virtual keys, which can be revoked.
- Quotas and budgets — requests per minute, tokens per minute and monthly dollar budgets per team or tenant.
- Routing — aliases map to one or more deployments; load balance between them.
- Reliability — retries with backoff, cooldown of failing deployments, fallback to another model.
- Caching — exact-match response cache; pass-through of provider prompt caching.
- Accounting — tokens and cost per team, feature and model, for chargeback.
- Observability and policy — one place to emit traces and run guardrails.
Tools
LiteLLM Proxy (open source) is common; others include Portkey, Kong AI Gateway, Cloudflare AI Gateway, and cloud-native options. A LiteLLM-style config looks like this:
1model_list:2 - model_name: chat-default # the alias apps call3 litellm_params:4 model: anthropic/claude-sonnet-4-55 api_key: os.environ/ANTHROPIC_API_KEY6 - model_name: chat-default # same alias, second provider7 litellm_params:8 model: bedrock/anthropic.claude-sonnet-4-5-20250929-v1:09 - model_name: chat-small # self-hosted vLLM10 litellm_params:11 model: hosted_vllm/meta-llama/Llama-3.1-8B-Instruct12 api_base: http://vllm.internal:8000/v11314router_settings:15 num_retries: 216 timeout: 30 # seconds per attempt17 allowed_fails: 3 # failures before a deployment cools down18 cooldown_time: 6019 fallbacks: [{"chat-default": ["chat-small"]}]Two deployments share the chat-default alias, so the gateway spreads load across the direct API and Bedrock. A deployment that fails three times is taken out for 60 seconds. If every chat-default deployment fails, the request falls back to the self-hosted small model.
The trade-offs
- An extra hop — a few milliseconds, small next to seconds of generation.
- A new single point of failure — run several replicas, keep it stateless (state in Redis), and health-check it.
- Buffering bugs — the gateway must pass streamed tokens through immediately.
A real-life example
A company's internal code assistant is used by 2,000 engineers across 40 teams, and other teams start building their own AI tools. Before the gateway, six services each held a provider key, and a leaked key in a CI log cost a weekend of rotation.
After the gateway: each team gets a virtual key with a monthly budget (say $1,500 for most teams, more for the assistant itself). When one team's nightly script loops and burns $900 in an hour, the gateway's budget stops it at the limit and alerts the owner. When the provider has an outage, the fallback to the self-hosted model keeps the assistant answering simple questions. And when the platform team moves to a newer model, they change one alias in one file.
Follow-up questions to expect
- "Where do you enforce rate limits — gateway or app?" — At the gateway, keyed by team or tenant, with shared state in Redis so limits hold across replicas.
- "Does the gateway break prompt caching?" — No, as long as it forwards the prompt unchanged; provider caches key on the prompt prefix, not on the client.
- "Build or buy?" — Start with an open-source proxy; build only the pieces that are specific to you, such as custom guardrails.