Course Content
Applied AI Engineering: From Prompt to Production
9 sections · 29 lessons
Context, cost and latency: the numbers to carry in your head
In the second week of the project, Harbourline's finance partner sent one line: "What will this cost per month, and what happens if usage doubles?" The team's first answer was "it depends". That is true and useless. The right answer takes ten minutes of arithmetic, and you should be able to do it before you write a line of retrieval code.
The same is true for latency. Before building, you can know roughly how long a question will take and which part of that time you can shrink. Engineers who carry a few numbers in their heads make better design choices in meetings, because they can reject a bad idea in thirty seconds instead of after a sprint.
This lesson gives you those numbers for PolicyPal, the code that measures the real ones, and a hard guard so a single bad request cannot blow the budget.
Token arithmetic for one question
Start with what goes into one call. PolicyPal's design, which Sections 2 and 3 build, sends a system prompt, a few retrieved policy chunks, a little conversation history and the question. Here is the budget.
| Part | Tokens | Notes |
|---|---|---|
| System prompt and rules | 700 | Fixed; the same on every call |
| 5 retrieved chunks at about 420 tokens | 2,100 | The largest part, and the one you control most |
| Last two turns of history | 150 | Trimmed; older turns are dropped |
| The question and user facts (country, grade) | 50 | Small |
| Input total | 3,000 | |
| Answer, including citations | 250 | Capped with max_tokens |
Writing the budget down forces useful questions. Why five chunks and not ten? Because ten would double the largest line and, as Section 3 shows, rarely finds the answer that five missed. Why only two turns of history? Because older turns cost tokens on every call and rarely matter for a policy question.
From tokens to money
Cost is a simple formula: input tokens times input price, plus output tokens times output price. With the mid-tier model at $2 per million input tokens and $10 per million output tokens:
- Input: 3,000 × $2 ÷ 1,000,000 = $0.0060
- Output: 250 × $10 ÷ 1,000,000 = $0.0025
- One question: about $0.0085
Harbourline expects around 1,200 questions per working day, and a month has about 22 working days. That is 26,400 questions, or about $224 a month. If usage doubles, it is about $450. That was the whole answer finance needed, and it took less time than the meeting.
Now compare the naive design. The 400 policy PDFs hold about 2.6 million tokens. That does not fit in any model's context window. Even the 45,000-token Version 0 prompt cost $0.09 per question before output, ten times the RAG design, and it covered only twelve policies. The arithmetic alone kills "just put everything in the prompt" for this problem.
A useful comparison is the human cost. An HR executive spends about 12 minutes on a routine policy ticket. At a loaded cost of about ₹450 an hour, that is ₹90, roughly $1. Even if PolicyPal only resolves half the questions it receives, the model bill is a rounding error next to the time returned to the helpdesk. The cost that deserves attention is the cost of a wrong answer, and no pricing page shows that.
From tokens to milliseconds
Latency adds up the same way. These are PolicyPal's measured medians once Section 3 is done.
| Step | Median time |
|---|---|
| Authentication and loading user profile | 10 ms |
| Router (Section 5) | 15 ms |
| Embedding the question and hybrid search | 30 ms |
| Reranking 40 candidates | 150 ms |
| Model time to first token (3,000-token prompt) | 700 ms |
| Model writing 250 tokens at about 70 tokens per second | 3,600 ms |
| Total | about 4.5 s |
Two things jump out. The model's writing time is 80% of the total, so the biggest lever is a shorter answer. And because the first token arrives at about 0.9 seconds, streaming the answer (Section 7) makes the experience feel fast even though the total barely changes. Optimising the 30-millisecond search step first would be wasted effort.
Count and log every call
Estimates are for planning. Once the system runs, you need real numbers, and every provider returns token counts with each response. PolicyPal records them on every call, tagged with the feature that made the call.
1# policypal/usage.py2import json3import logging45from policypal.llm import Reply67# USD per million tokens (input, output). List prices at the time of writing.8PRICES = {9 "claude-opus-5": (5.00, 25.00),10 "claude-sonnet-5": (2.00, 10.00),11 "claude-haiku-4-5": (1.00, 5.00),12}13log = logging.getLogger("policypal.usage")1415class ContextBudgetExceeded(Exception):16 pass1718def estimate_tokens(text: str) -> int:19 return len(text) // 4 + 1 # rough, English-biased; fine for a guard2021def check_budget(system: str, messages: list[dict], limit: int = 8_000) -> None:22 text = system + "".join(str(m["content"]) for m in messages)23 if estimate_tokens(text) > limit:24 raise ContextBudgetExceeded(f"about {estimate_tokens(text)} tokens, limit {limit}")2526def record(model: str, feature: str, reply: Reply) -> float:27 price_in, price_out = PRICES[model]28 usd = (reply.input_tokens * price_in + reply.output_tokens * price_out) / 1_000_00029 log.info(json.dumps({"model": model, "feature": feature,30 "in": reply.input_tokens, "out": reply.output_tokens,31 "ms": reply.latency_ms, "usd": round(usd, 6)}))32 return usdrecord writes one JSON line per call, so a daily query over the logs gives cost per feature, per model and per day. The feature tag matters more than it looks: when the bill jumps, you want to know within a minute whether it was answers, the agent loop or a nightly eval job.
check_budget is a cheap guard, not an exact count. It stops a request that is far over budget, such as a user pasting a 60-page document, before it reaches the model. The limit of 8,000 is more than twice the normal 3,000, so it never blocks ordinary traffic. When you need an exact count, providers offer a token-counting endpoint; for the Anthropic SDK it is client.messages.count_tokens(...). It costs an extra network round trip, so use it where precision matters, not on every request.
The numbers to carry
These are rules of thumb, not laws. They are close enough to make design decisions in a meeting.
- 1 token is about 0.75 English words; one policy page is about 600 tokens.
- Output tokens cost four to five times as much as input tokens.
- Time to first token is about 0.5 to 1 second for a prompt of a few thousand tokens.
- Output speed is about 50 to 100 tokens per second on hosted models.
- Cached input tokens (Section 7) cost about a tenth of normal input tokens.
- A retrieval step over a few thousand chunks takes tens of milliseconds; a reranker takes a hundred or more on a CPU.
Check your understanding
0 of 3 answered
1.PolicyPal's answers grow from 250 to 500 tokens after a prompt change, with input unchanged at 3,000 tokens. Roughly how does cost per question change on the mid-tier model?
2.Which step should the team optimise first to make PolicyPal feel faster?
3.Why does check_budget use a rough estimate of four characters per token instead of the provider's exact counting endpoint?