Course Content
LLMOps & Deployment
6 sections · 40 lessons
What techniques can reduce inference costs without hurting quality?
What you need to know
Levers ranked by safety
| Lever | What it does | Risk to quality |
|---|---|---|
| Prompt caching | Reuses the processed static prefix; cached input is heavily discounted | None — same output |
| Exact-match cache | Returns a stored answer for an identical normalised request | None, if keyed on prompt and model version |
Cap max_tokens, ask for concise answers | Fewer output tokens, the expensive kind | Low; check truncation rate |
| Batch API | ~50% off for work that can wait hours | None; only latency |
| Trim the prompt | Fewer, better chunks; summary instead of full history | Medium — eval it |
| Model routing | Easy requests go to a small model | Medium — eval it |
| Semantic cache | Reuses an answer for a similar question | High — can return a wrong answer |
| Self-host + quantize | Lower cost per token at high utilisation | Low to medium — eval it |
How prompt caching works
The model processes the prompt from left to right. If the first 4,000 tokens are the same as a recent request, the provider (or vLLM's prefix cache) reuses the saved KV cache for them. You pay much less for those tokens, and time to first token falls too. The rule: static content first (system prompt, tool definitions, examples, long documents), variable content last (the user's question). A timestamp at the top of the prompt breaks the cache for everything after it.
Why semantic caching is risky
A semantic cache embeds the question and returns a stored answer if a past question is "close enough". But "How do I cancel my order?" and "How do I cancel my subscription?" are very close in embedding space and need different answers. Use it only for FAQ-like traffic, with a strict similarity threshold, and scope the cache by user or tenant when answers depend on account data.
A real-life example
An internal code assistant serves 2,000 engineers, about 60 requests each per working day: 2.64 million requests a month. Each sends 6,000 input tokens (instructions, tool definitions, repository context) and gets 400 output tokens. At illustrative prices of $3/$15 per million, that is $0.024 per request, or $63,000 a month.
The team applies levers in order:
- Prompt caching. They reorder the prompt so 4,000 tokens of instructions and shared repository context come first. Cost per request falls to $0.0132: $34,800 a month.
- Exact-match cache. 10% of requests repeat exactly (the same "explain this error" on a popular build failure): $31,400.
- Routing. A small classifier sends 60% of requests (explain, rename, write a docstring) to a small model that costs about a tenth as much per request. The eval set shows no drop on those task types but a clear drop on refactoring, so refactoring stays on the large model: about $14,500 a month.
Cost falls by 77%, and every step after the first went through the eval gate first.
Follow-up questions to expect
- "Which lever would you try first?" — Prompt caching and output length: they are cheap to do and do not change answers.
- "How do you prove quality did not drop?" — Run the eval set before and after, per task type, and watch live signals like thumbs-down and regenerate rate during a canary.
- "Does a bigger context window mean I should send more?" — No. Longer prompts cost more, are slower, and can reduce accuracy; retrieve less but better.