Applied AI Engineering: From Prompt to Production

Course Content

Applied AI Engineering: From Prompt to Production

9 sections · 29 lessons

Context, cost and latency: the numbers to carry in your head


In the second week of the project, Harbourline's finance partner sent one line: "What will this cost per month, and what happens if usage doubles?" The team's first answer was "it depends". That is true and useless. The right answer takes ten minutes of arithmetic, and you should be able to do it before you write a line of retrieval code.

The same is true for latency. Before building, you can know roughly how long a question will take and which part of that time you can shrink. Engineers who carry a few numbers in their heads make better design choices in meetings, because they can reject a bad idea in thirty seconds instead of after a sprint.

This lesson gives you those numbers for PolicyPal, the code that measures the real ones, and a hard guard so a single bad request cannot blow the budget.

Where 4.5 seconds go, in milliseconds1015301507003600012345authsearchrerankfirst tokenwriting250 tokensAbout 3,000 input and 250 output tokens: roughly 0.0085 dollars per question.
The model writing its answer is 80% of the time, so shorter answers and streaming beat any search optimisation.

Token arithmetic for one question

Start with what goes into one call. PolicyPal's design, which Sections 2 and 3 build, sends a system prompt, a few retrieved policy chunks, a little conversation history and the question. Here is the budget.

PartTokensNotes
System prompt and rules700Fixed; the same on every call
5 retrieved chunks at about 420 tokens2,100The largest part, and the one you control most
Last two turns of history150Trimmed; older turns are dropped
The question and user facts (country, grade)50Small
Input total3,000
Answer, including citations250Capped with max_tokens

Writing the budget down forces useful questions. Why five chunks and not ten? Because ten would double the largest line and, as Section 3 shows, rarely finds the answer that five missed. Why only two turns of history? Because older turns cost tokens on every call and rarely matter for a policy question.

From tokens to money

Cost is a simple formula: input tokens times input price, plus output tokens times output price. With the mid-tier model at $2 per million input tokens and $10 per million output tokens:

  • Input: 3,000 × $2 ÷ 1,000,000 = $0.0060
  • Output: 250 × $10 ÷ 1,000,000 = $0.0025
  • One question: about $0.0085

Harbourline expects around 1,200 questions per working day, and a month has about 22 working days. That is 26,400 questions, or about $224 a month. If usage doubles, it is about $450. That was the whole answer finance needed, and it took less time than the meeting.

Now compare the naive design. The 400 policy PDFs hold about 2.6 million tokens. That does not fit in any model's context window. Even the 45,000-token Version 0 prompt cost $0.09 per question before output, ten times the RAG design, and it covered only twelve policies. The arithmetic alone kills "just put everything in the prompt" for this problem.

A useful comparison is the human cost. An HR executive spends about 12 minutes on a routine policy ticket. At a loaded cost of about ₹450 an hour, that is ₹90, roughly $1. Even if PolicyPal only resolves half the questions it receives, the model bill is a rounding error next to the time returned to the helpdesk. The cost that deserves attention is the cost of a wrong answer, and no pricing page shows that.

From tokens to milliseconds

Latency adds up the same way. These are PolicyPal's measured medians once Section 3 is done.

StepMedian time
Authentication and loading user profile10 ms
Router (Section 5)15 ms
Embedding the question and hybrid search30 ms
Reranking 40 candidates150 ms
Model time to first token (3,000-token prompt)700 ms
Model writing 250 tokens at about 70 tokens per second3,600 ms
Totalabout 4.5 s

Two things jump out. The model's writing time is 80% of the total, so the biggest lever is a shorter answer. And because the first token arrives at about 0.9 seconds, streaming the answer (Section 7) makes the experience feel fast even though the total barely changes. Optimising the 30-millisecond search step first would be wasted effort.

Count and log every call

Estimates are for planning. Once the system runs, you need real numbers, and every provider returns token counts with each response. PolicyPal records them on every call, tagged with the feature that made the call.

Python
# policypal/usage.pyimport jsonimport loggingfrom policypal.llm import Reply# USD per million tokens (input, output). List prices at the time of writing.PRICES = {    "claude-opus-5": (5.00, 25.00),    "claude-sonnet-5": (2.00, 10.00),    "claude-haiku-4-5": (1.00, 5.00),}log = logging.getLogger("policypal.usage")class ContextBudgetExceeded(Exception):    passdef estimate_tokens(text: str) -> int:    return len(text) // 4 + 1          # rough, English-biased; fine for a guarddef check_budget(system: str, messages: list[dict], limit: int = 8_000) -> None:    text = system + "".join(str(m["content"]) for m in messages)    if estimate_tokens(text) > limit:        raise ContextBudgetExceeded(f"about {estimate_tokens(text)} tokens, limit {limit}")def record(model: str, feature: str, reply: Reply) -> float:    price_in, price_out = PRICES[model]    usd = (reply.input_tokens * price_in + reply.output_tokens * price_out) / 1_000_000    log.info(json.dumps({"model": model, "feature": feature,                         "in": reply.input_tokens, "out": reply.output_tokens,                         "ms": reply.latency_ms, "usd": round(usd, 6)}))    return usd

record writes one JSON line per call, so a daily query over the logs gives cost per feature, per model and per day. The feature tag matters more than it looks: when the bill jumps, you want to know within a minute whether it was answers, the agent loop or a nightly eval job.

check_budget is a cheap guard, not an exact count. It stops a request that is far over budget, such as a user pasting a 60-page document, before it reaches the model. The limit of 8,000 is more than twice the normal 3,000, so it never blocks ordinary traffic. When you need an exact count, providers offer a token-counting endpoint; for the Anthropic SDK it is client.messages.count_tokens(...). It costs an extra network round trip, so use it where precision matters, not on every request.

The numbers to carry

These are rules of thumb, not laws. They are close enough to make design decisions in a meeting.

  • 1 token is about 0.75 English words; one policy page is about 600 tokens.
  • Output tokens cost four to five times as much as input tokens.
  • Time to first token is about 0.5 to 1 second for a prompt of a few thousand tokens.
  • Output speed is about 50 to 100 tokens per second on hosted models.
  • Cached input tokens (Section 7) cost about a tenth of normal input tokens.
  • A retrieval step over a few thousand chunks takes tens of milliseconds; a reranker takes a hundred or more on a CPU.

Check your understanding

0 of 3 answered

1.PolicyPal's answers grow from 250 to 500 tokens after a prompt change, with input unchanged at 3,000 tokens. Roughly how does cost per question change on the mid-tier model?

2.Which step should the team optimise first to make PolicyPal feel faster?

3.Why does check_budget use a rough estimate of four characters per token instead of the provider's exact counting endpoint?