LLMOps & Deployment

Course Content

LLMOps & Deployment

6 sections · 40 lessons

What techniques can reduce inference costs without hurting quality?


Code assistant monthly bill, lever by lever (thousands)63.434.831.414.50123baselineprompt cachingexact-matchcacheroute 60%to small2.64 million requests a month; only the last step changed answers, so only it needed the eval gate.
The biggest safe saving came from reordering the prompt, before any change that could touch quality.

What you need to know

Levers ranked by safety

LeverWhat it doesRisk to quality
Prompt cachingReuses the processed static prefix; cached input is heavily discountedNone — same output
Exact-match cacheReturns a stored answer for an identical normalised requestNone, if keyed on prompt and model version
Cap max_tokens, ask for concise answersFewer output tokens, the expensive kindLow; check truncation rate
Batch API~50% off for work that can wait hoursNone; only latency
Trim the promptFewer, better chunks; summary instead of full historyMedium — eval it
Model routingEasy requests go to a small modelMedium — eval it
Semantic cacheReuses an answer for a similar questionHigh — can return a wrong answer
Self-host + quantizeLower cost per token at high utilisationLow to medium — eval it

How prompt caching works

The model processes the prompt from left to right. If the first 4,000 tokens are the same as a recent request, the provider (or vLLM's prefix cache) reuses the saved KV cache for them. You pay much less for those tokens, and time to first token falls too. The rule: static content first (system prompt, tool definitions, examples, long documents), variable content last (the user's question). A timestamp at the top of the prompt breaks the cache for everything after it.

Why semantic caching is risky

A semantic cache embeds the question and returns a stored answer if a past question is "close enough". But "How do I cancel my order?" and "How do I cancel my subscription?" are very close in embedding space and need different answers. Use it only for FAQ-like traffic, with a strict similarity threshold, and scope the cache by user or tenant when answers depend on account data.

A real-life example

An internal code assistant serves 2,000 engineers, about 60 requests each per working day: 2.64 million requests a month. Each sends 6,000 input tokens (instructions, tool definitions, repository context) and gets 400 output tokens. At illustrative prices of $3/$15 per million, that is $0.024 per request, or $63,000 a month.

The team applies levers in order:

  1. Prompt caching. They reorder the prompt so 4,000 tokens of instructions and shared repository context come first. Cost per request falls to $0.0132: $34,800 a month.
  2. Exact-match cache. 10% of requests repeat exactly (the same "explain this error" on a popular build failure): $31,400.
  3. Routing. A small classifier sends 60% of requests (explain, rename, write a docstring) to a small model that costs about a tenth as much per request. The eval set shows no drop on those task types but a clear drop on refactoring, so refactoring stays on the large model: about $14,500 a month.

Cost falls by 77%, and every step after the first went through the eval gate first.

Follow-up questions to expect

  • "Which lever would you try first?" — Prompt caching and output length: they are cheap to do and do not change answers.
  • "How do you prove quality did not drop?" — Run the eval set before and after, per task type, and watch live signals like thumbs-down and regenerate rate during a canary.
  • "Does a bigger context window mean I should send more?" — No. Longer prompts cost more, are slower, and can reduce accuracy; retrieve less but better.