Course Content
Applied AI Engineering: From Prompt to Production
9 sections · 29 lessons
Latency and throughput: caching, batching, streaming
After launch, the most common piece of feedback was not about accuracy. It was "It's a bit slow." The median question took 4.5 seconds, and for all of that time the page showed a spinner. On announcement days it was worse: during the hybrid-work policy email, p95 latency reached 11 seconds, because the provider's rate limit forced retries.
You already know where the time goes from Section 1: about 80% of it is the model writing the answer. No amount of tuning the 30-millisecond search step will change how the product feels. What changes it is showing words as they are written, not repeating work you have already done, and moving work that does not need to be instant out of the request path.
This lesson covers the three tools for that: streaming, caching and batching, each with its trade-off made explicit.
Streaming a structured answer
Streaming sends the answer to the browser token by token, so the first words appear at the time to first token, about 0.9 seconds, instead of after the whole answer. PolicyPal's complication is that its reply is JSON, and a half-written JSON object is not something you can show.
The trick is to parse the partial JSON on every chunk and forward only the new characters of the answer field. The jiter library, which the Anthropic SDK already depends on, can parse incomplete JSON and return partial strings. The LLM client gains a streaming method that wraps the SDK's messages.stream(...) and yields text chunks.
1# policypal/streaming.py2import json34import jiter56from policypal.schemas import ANSWER_SCHEMA78def sse(data, event: str | None = None) -> str:9 head = f"event: {event}\n" if event else ""10 return f"{head}data: {json.dumps(data)}\n\n"1112def stream_answer(pipeline, user, question: str):13 prepared = pipeline.prepare(user, question) # route and retrieve: about 200 ms14 buf, sent, status_sent = "", 0, False15 for chunk in pipeline.llm.stream_text(prepared.system, prepared.messages,16 schema=ANSWER_SCHEMA, max_tokens=600):17 buf += chunk18 try:19 partial = jiter.from_json(buf.encode(), partial_mode="trailing-strings")20 except ValueError:21 continue # not parseable yet22 if not isinstance(partial, dict):23 continue24 if not status_sent and "answer" in partial: # status is complete once answer starts25 yield sse(partial["status"], event="status")26 status_sent = True27 text = partial.get("answer", "")28 if len(text) > sent:29 yield sse(text[sent:])30 sent = len(text)31 final = pipeline.finish(user, prepared, buf) # validate, ground-check, log, trace32 yield sse(final.model_dump(), event="done")The FastAPI route returns StreamingResponse(stream_answer(...), media_type="text/event-stream"). Because status comes first in the schema, as planned in Section 2, the page knows whether to show "Talk to HR" before the answer text starts. Re-parsing the whole buffer on each chunk sounds wasteful, but an answer is under 2,000 characters, so it costs microseconds.
Streaming has an honest cost: the grounding check can only run on the complete answer, after the user has already read it. PolicyPal handles this in the done event. If the check dropped a sentence, the page replaces the streamed text with the checked version and shows "updated after checking sources". That happens on about 5% of answers. For the few routes where even a brief unchecked sentence is unacceptable, such as answers involving money above a threshold, PolicyPal does not stream; it shows a progress message and the checked answer.
With streaming, median time to the first visible word fell to 0.9 seconds, and p95 to 1.8. The total time barely changed, and the "slow" feedback nearly stopped.
Provider prompt caching
Providers can cache the processed form of a prompt's prefix, so later requests that start with exactly the same tokens are cheaper and faster. Cached input tokens typically cost about a tenth of normal input tokens. Two conditions apply: the prefix must be byte-for-byte identical, and it must be longer than a model-specific minimum, often one to four thousand tokens.
That second condition matters for PolicyPal. The answer path's stable prefix is only the 700-token system prompt, because the examples and sources change per question; it is below the minimum on many models. The agent path is different. Its system prompt and tool definitions are about 2,000 tokens, and every step of the loop re-sends the whole conversation so far. A three-step run sends about 3,500, 5,000 and 6,500 input tokens, and most of each request is the previous request repeated.
With the Anthropic SDK, the simplest form is a top-level setting on the request. In PolicyPal's client it is one line in complete:
kwargs["cache_control"] = {"type": "ephemeral"} # cache the longest reusable prefixCheck that it works with resp.usage.cache_read_input_tokens, which PolicyPal now logs next to input and output tokens. If it stays at zero, something in the prefix is changing between calls: a timestamp in the system prompt, tool definitions in a different order, or a request id. On the agent path, caching cut input cost by about 30%.
Caching answers, carefully
An answer cache skips the whole pipeline when the same question has been answered before. On a normal day only about 9% of policy questions repeat exactly after normalising case and spacing. On announcement days, it was 31%, which is exactly when it helps most.
The danger is serving an answer to someone it does not apply to. So PolicyPal's cache key includes everything that changes the answer, and it caches only answers that are the same for everyone with that key.
1# policypal/answer_cache.py2import hashlib3import re45import redis67r = redis.Redis(host="cache", port=6379)89def cache_key(question: str, user, release: str) -> str:10 norm = re.sub(r"\s+", " ", question.lower()).strip(" ?!.")11 raw = f"{release}|{user.country}|{user.grade}|{norm}"12 return "ans:" + hashlib.sha256(raw.encode()).hexdigest()1314def get_cached(question: str, user, release: str) -> str | None:15 value = r.get(cache_key(question, user, release))16 return value.decode() if value else None1718def put_cached(question: str, user, release: str, route: str, answer_json: str, status: str) -> None:19 if route == "policy_question" and status == "answered": # never personal or tool results20 r.set(cache_key(question, user, release), answer_json, ex=24 * 3600)The release id in the key means a new prompt, model or index automatically stops old answers being served. Country and grade are in the key because policies differ by both. Only policy_question answers with status answered are stored: a leave balance, a ticket or a handover is personal and must never be served to someone else. The 24-hour expiry is a backstop.
A semantic cache, which returns a stored answer for a similar question, is tempting and dangerous here. "Can I carry forward 10 days?" and "Can I carry forward 40 days?" are 97% similar by embedding and have opposite answers, because the limit is 30. PolicyPal does not use one. It would reconsider only with a verification step that checks the stored answer against the new question, which removes most of the saving.
Batching the work that can wait
Not everything needs to be instant. The nightly index build embeds about 6,200 chunks; encoding them in batches of 64 takes about 2.5 minutes on a CPU, against about 12 minutes one at a time, because the model processes a batch in one pass. The eval suite does not need answers in seconds either. Providers offer batch APIs for exactly this: on the Anthropic API, client.messages.batches.create(...) accepts many requests at once, returns results within hours, and charges about half the normal price. PolicyPal's nightly eval run, 250 cases plus 60 agent cases three times, uses it.
Batching also helps with bursts, from the other direction. HR now tells the PolicyPal team the day before a policy email goes out. The team writes the 20 most likely questions, runs them through the pipeline in advance, and the answer cache is warm before the first employee opens the email.
Check your understanding
0 of 3 answered
1.Why does PolicyPal include the release id in the answer cache key?
2.A teammate proposes a semantic cache that returns a stored answer when a new question is at least 95% similar by embedding. What is the main risk for PolicyPal?
3.PolicyPal streams answers, but the NLI grounding check needs the complete answer. How does it handle this?