LLMOps & Deployment

Course Content

LLMOps & Deployment

6 sections · 40 lessons

How do you estimate the cost of running an LLM-powered feature?


What one support-bot call actually billsAnswer(output): 500 tokensConversationhistory: 300Retrievedchunks: 1,500System prompt +examples: 1,200topbottomAt 3 and 15 per million: 0.009 for 3,000 input tokens, 0.0075 for 500 output tokens.
Output is one seventh of the tokens but almost half the bill, which is why answer length is the first lever to pull.

What you need to know

The base formula

Text
cost_per_call = (input_tokens  / 1,000,000) × input_price              + (output_tokens / 1,000,000) × output_price

Prices below are illustrative ($3 per million input tokens, $15 per million output); always use your provider's current price sheet.

Input tokens include everything the model reads: system prompt, few-shot examples, retrieved chunks, conversation history, tool results. Output tokens are what it writes — and on reasoning models, the hidden reasoning tokens are billed as output too.

Worked example

A support bot sends a 3,000-token prompt (1,200 system + examples, 1,500 retrieved chunks, 300 history) and gets a 500-token answer:

Text
input:  3,000 / 1M × $3  = $0.0090output:   500 / 1M × $15 = $0.0075total per call           = $0.0165

The multipliers people forget

  • Calls per user turn. An agent that plans, calls a tool and then answers makes 3 calls, so 3× the cost.
  • Retries, guardrail and judge calls — often 10–30% extra.
  • Growing history. Turn 10 of a chat carries turns 1–9 as input.
  • Embeddings, reranking and vector DB hosting.

The discounts

  • Prompt caching. Put the static part first. Major providers bill cached input at a large discount — up to about 90% off on current models (illustrative; check the sheet). Cache writes can cost slightly more.
  • Batch APIs for offline jobs — typically about 50% off, with results within 24 hours.
  • Routing easy requests to a cheaper model.

A small calculator makes the assumptions visible:

Python
PRICE_IN, PRICE_OUT = 3.00, 15.00       # $ per 1M tokens (illustrative)def monthly_cost(requests, inp, out, calls_per_request=1,                 cached_in=0, cache_discount=0.9, overhead=0.15):    fresh_in = inp - cached_in    per_call = (fresh_in * PRICE_IN                + cached_in * PRICE_IN * (1 - cache_discount)                + out * PRICE_OUT) / 1e6    return requests * calls_per_request * per_call * (1 + overhead)normal = monthly_cost(requests=4_800_000, inp=3000, out=500)cached = monthly_cost(requests=4_800_000, inp=3000, out=500, cached_in=2000)print(f"no caching:   ${normal:,.0f}/month")print(f"with caching: ${cached:,.0f}/month")

It prints $91,080 without caching and $61,272 with 2,000 of the 3,000 input tokens cached. overhead=0.15 covers retries and guardrail calls.

A real-life example

A fintech's support bot handles 40,000 conversations a day, 4 turns each: 160,000 calls a day, or 4.8 million a month. At $0.0165 per call that is about $2,640 a day, and $91,000 a month with 15% overhead.

Then the Diwali sale arrives. Traffic rises 5× for a week: $13,200 a day in model cost alone. The finance team asks for a forecast before the sale, so the engineer models three levers: caching the 2,000-token static prefix (cost falls a third), capping answers at 350 tokens instead of 500, and sending order-status questions — 40% of traffic — to a small model. The forecast goes to finance with a per-day budget alert set at $9,000.

Follow-up questions to expect

  • "Why are output tokens more expensive?" — Input tokens are processed in parallel in one pass (prefill); output tokens are generated one at a time, each needing a full pass over the model weights.
  • "How do you estimate before launch with no traffic?" — Run the eval set, log real token counts per case, and multiply by expected volume with a safety margin.
  • "How does self-hosting change the model?" — You pay for GPU hours whether used or not, so cost per token depends on utilisation, not token count.