Course Content
Applied AI Engineering: From Prompt to Production
9 sections · 29 lessons
Serving the app as a service
Until now, PolicyPal has been a set of Python modules, an eval harness and a demo page that one engineer ran from a laptop. The pilot group of 200 people forgave the occasional restart. The rollout to all 2,000 employees would not.
A service has jobs a script never has. It must check who is asking, on every request. It must load a 130 MB embedding model and a reranker once, not per question. It must survive the model provider being slow for twenty minutes, the HR system being down for maintenance, and a policy announcement that brings 400 questions in an hour. And when someone reports a bad answer three weeks later, it must be able to say exactly which prompt, model and index produced it.
This lesson turns PolicyPal into that service. Nothing here is specific to AI except the last part, and the last part is the one teams most often get wrong.
The shape of the service
PolicyPal runs as a small set of processes, each placed where its resources fit.
| Component | Where it runs | Why there |
|---|---|---|
| API (FastAPI) with retrieval, reranker, NLI check | 3 replicas, 2 vCPU and 4 GB each | CPU models, about 1.2 GB loaded per replica |
| Router (0.5B LoRA model) | One small GPU service, shared | 15 ms on GPU against 60 on CPU |
| Answering and judge models | Hosted provider, via the LLM client | No GPUs to own; pay per token |
| Index build | Nightly job | Heavy, and must not slow live traffic |
| Answer cache | Redis | Shared across replicas |
The router sits behind its own small HTTP API with a 100-millisecond timeout. If it does not answer in time, the API falls back to the prompted router from Section 4, which is slower but always available. That single decision means a GPU problem makes PolicyPal slower, never unavailable.
The request path in code
1# policypal/app.py2import asyncio3from contextlib import asynccontextmanager45from fastapi import Depends, FastAPI, HTTPException6from pydantic import BaseModel, Field78from policypal.auth import User, current_user # validates the company SSO token9from policypal.config import settings10from policypal.pipeline import Pipeline # route, retrieve, answer, check1112class AskIn(BaseModel):13 question: str = Field(min_length=1, max_length=4000)14 conversation_id: str | None = None1516@asynccontextmanager17async def lifespan(app: FastAPI):18 app.state.pipeline = Pipeline.load(settings) # embedder, BM25, reranker, NLI19 yield2021app = FastAPI(lifespan=lifespan)2223@app.get("/healthz")24def healthz() -> dict:25 return {"ok": True} # the process is alive2627@app.get("/readyz")28def readyz() -> dict:29 if not hasattr(app.state, "pipeline"):30 raise HTTPException(503, "loading") # do not send traffic yet31 return {"release": settings.release}3233@app.post("/ask")34async def ask(body: AskIn, user: User = Depends(current_user)) -> dict:35 pipeline = app.state.pipeline36 try:37 result = await asyncio.wait_for(38 asyncio.to_thread(pipeline.answer, user, body.question, body.conversation_id),39 timeout=settings.request_timeout_s) # 25 s40 except asyncio.TimeoutError:41 raise HTTPException(504, "PolicyPal took too long. Please try again.")42 return result.model_dump() | {"release": settings.release}Models load once, in the lifespan function, before the service reports itself ready. Loading takes about 40 seconds, which is why there are two health endpoints: /healthz says the process is alive, and /readyz says it can take traffic. The load balancer only sends requests to replicas that are ready.
The user comes from a dependency that validates the company's single sign-on token. Everything downstream, the country filter, the leave tool and the audit log, uses that User, never anything from the request body. The question has a hard length limit, enforced by Pydantic before any model sees it.
pipeline.answer is ordinary blocking code: the reranker and NLI model are CPU-bound, and the model client is synchronous. Running it with asyncio.to_thread keeps the event loop free for other requests. The 25-second timeout gives the user a clear message instead of a hanging page. One caution: the timeout stops waiting, but the thread keeps running until it finishes, so the model client's own 30-second timeout from Section 1 is what finally bounds the work.
Capacity from real numbers
How many requests does PolicyPal handle at once? Little's law gives the answer: the average number of requests in flight equals the arrival rate times the time each one takes. On a normal peak, 30 questions a minute is 0.5 per second, and each takes about 4.5 seconds, so about 2 to 3 requests are in flight. Three replicas are about resilience, not load.
The day that matters is an announcement. When HR emailed all staff about a new hybrid-work policy, questions arrived at 200 a minute for half an hour: about 3.3 per second, or 15 requests in flight. The API replicas coped easily. The limit that bit was the provider's rate limit, counted in tokens per minute: 200 questions a minute at 3,000 input tokens each is 600,000 input tokens a minute, above the account's tier at the time. Requests started receiving 429 errors, retries piled up, and latency climbed. The fix was part capacity (a higher tier, arranged in advance) and part design (the caching in the next lesson).
Degraded modes, decided in advance
Every dependency will fail at some point. Decide what PolicyPal does in each case before it happens, and test each mode.
| Failure | Degraded behaviour | Cost to users |
|---|---|---|
| Router service down | Prompted router | About 850 ms slower |
| Reranker too slow | Skip reranking, use hybrid top 5 | Recall@5 falls from 0.91 to 0.84 |
| Answering model errors or times out | Retry once, then a second, smaller model | Slightly lower quality |
| Provider fully unavailable | Search-only mode: show the top 3 policy sections with links | No written answer |
| HRMS down | Tool returns an error; the model explains | No balances for a while |
Search-only mode deserves a mention. When no model is available, PolicyPal still does something useful: it shows the three best-matching policy sections, reranked, with a note that answers are temporarily unavailable. Employees used it twice in the first quarter and barely noticed.
A release is more than code
A PolicyPal answer depends on six things: the code, the answering model, the prompt, the example bank, the router adapter and the index. Change any one and answers change. So a release pins all six in a manifest, and every answer is logged with the release id.
1# releases/2026.03.18-2.yaml2release: 2026.03.18-23code: git 4f2c9ab4answer_model: claude-sonnet-55fallback_model: claude-haiku-4-56judge_model: claude-opus-57prompt: answer-v2.48examples: bank-v79router_adapter: router-2026-03-0210index: policies-2026-03-18T02-00Z # 6,214 chunks from 397 documentsThe nightly index build writes a new, immutable index directory with a timestamped name. It is not used until a release manifest points to it, after the retrieval metrics from Section 6 pass on the new index. Switching releases is a configuration change that replicas pick up on restart; rolling back is pointing at the previous manifest. Nothing is edited in place.
This is the difference between "the answers got worse this week" and "release 2026.03.18-2 changed the example bank from v6 to v7; here are the 11 eval cases it broke".
Check your understanding
0 of 3 answered
1.At a normal peak of 30 questions a minute, each taking 4.5 seconds, about how many requests are in flight at once?
2.Why does PolicyPal have both /healthz and /readyz endpoints?
3.An answer from three weeks ago is reported as wrong. What makes it possible to find the cause quickly?