Course Content
LLMOps & Deployment
6 sections · 40 lessons
How do you estimate the cost of running an LLM-powered feature?
What you need to know
The base formula
cost_per_call = (input_tokens / 1,000,000) × input_price + (output_tokens / 1,000,000) × output_pricePrices below are illustrative ($3 per million input tokens, $15 per million output); always use your provider's current price sheet.
Input tokens include everything the model reads: system prompt, few-shot examples, retrieved chunks, conversation history, tool results. Output tokens are what it writes — and on reasoning models, the hidden reasoning tokens are billed as output too.
Worked example
A support bot sends a 3,000-token prompt (1,200 system + examples, 1,500 retrieved chunks, 300 history) and gets a 500-token answer:
input: 3,000 / 1M × $3 = $0.0090output: 500 / 1M × $15 = $0.0075total per call = $0.0165The multipliers people forget
- Calls per user turn. An agent that plans, calls a tool and then answers makes 3 calls, so 3× the cost.
- Retries, guardrail and judge calls — often 10–30% extra.
- Growing history. Turn 10 of a chat carries turns 1–9 as input.
- Embeddings, reranking and vector DB hosting.
The discounts
- Prompt caching. Put the static part first. Major providers bill cached input at a large discount — up to about 90% off on current models (illustrative; check the sheet). Cache writes can cost slightly more.
- Batch APIs for offline jobs — typically about 50% off, with results within 24 hours.
- Routing easy requests to a cheaper model.
A small calculator makes the assumptions visible:
1PRICE_IN, PRICE_OUT = 3.00, 15.00 # $ per 1M tokens (illustrative)23def monthly_cost(requests, inp, out, calls_per_request=1,4 cached_in=0, cache_discount=0.9, overhead=0.15):5 fresh_in = inp - cached_in6 per_call = (fresh_in * PRICE_IN7 + cached_in * PRICE_IN * (1 - cache_discount)8 + out * PRICE_OUT) / 1e69 return requests * calls_per_request * per_call * (1 + overhead)1011normal = monthly_cost(requests=4_800_000, inp=3000, out=500)12cached = monthly_cost(requests=4_800_000, inp=3000, out=500, cached_in=2000)13print(f"no caching: ${normal:,.0f}/month")14print(f"with caching: ${cached:,.0f}/month")It prints $91,080 without caching and $61,272 with 2,000 of the 3,000 input tokens cached. overhead=0.15 covers retries and guardrail calls.
A real-life example
A fintech's support bot handles 40,000 conversations a day, 4 turns each: 160,000 calls a day, or 4.8 million a month. At $0.0165 per call that is about $2,640 a day, and $91,000 a month with 15% overhead.
Then the Diwali sale arrives. Traffic rises 5× for a week: $13,200 a day in model cost alone. The finance team asks for a forecast before the sale, so the engineer models three levers: caching the 2,000-token static prefix (cost falls a third), capping answers at 350 tokens instead of 500, and sending order-status questions — 40% of traffic — to a small model. The forecast goes to finance with a per-day budget alert set at $9,000.
Follow-up questions to expect
- "Why are output tokens more expensive?" — Input tokens are processed in parallel in one pass (prefill); output tokens are generated one at a time, each needing a full pass over the model weights.
- "How do you estimate before launch with no traffic?" — Run the eval set, log real token counts per case, and multiply by expected volume with a safety margin.
- "How does self-hosting change the model?" — You pay for GPU hours whether used or not, so cost per token depends on utilisation, not token count.