Applied AI Engineering: From Prompt to Production

Course Content

Applied AI Engineering: From Prompt to Production

9 sections · 29 lessons

Serving the app as a service


Until now, PolicyPal has been a set of Python modules, an eval harness and a demo page that one engineer ran from a laptop. The pilot group of 200 people forgave the occasional restart. The rollout to all 2,000 employees would not.

A service has jobs a script never has. It must check who is asking, on every request. It must load a 130 MB embedding model and a reranker once, not per question. It must survive the model provider being slow for twenty minutes, the HR system being down for maintenance, and a policy announcement that brings 400 questions in an hour. And when someone reports a bad answer three weeks later, it must be able to say exactly which prompt, model and index produced it.

This lesson turns PolicyPal into that service. Nothing here is specific to AI except the last part, and the last part is the one teams most often get wrong.

What one release pinsRelease manifestCode commitAnswering andfallback modelsPrompt versionExample bankRouter adapterImmutable index snapshot
Six things shape every answer, so each answer logs its release and rollback is switching manifests.

The shape of the service

PolicyPal runs as a small set of processes, each placed where its resources fit.

ComponentWhere it runsWhy there
API (FastAPI) with retrieval, reranker, NLI check3 replicas, 2 vCPU and 4 GB eachCPU models, about 1.2 GB loaded per replica
Router (0.5B LoRA model)One small GPU service, shared15 ms on GPU against 60 on CPU
Answering and judge modelsHosted provider, via the LLM clientNo GPUs to own; pay per token
Index buildNightly jobHeavy, and must not slow live traffic
Answer cacheRedisShared across replicas

The router sits behind its own small HTTP API with a 100-millisecond timeout. If it does not answer in time, the API falls back to the prompted router from Section 4, which is slower but always available. That single decision means a GPU problem makes PolicyPal slower, never unavailable.

The request path in code

Python
# policypal/app.pyimport asynciofrom contextlib import asynccontextmanagerfrom fastapi import Depends, FastAPI, HTTPExceptionfrom pydantic import BaseModel, Fieldfrom policypal.auth import User, current_user      # validates the company SSO tokenfrom policypal.config import settingsfrom policypal.pipeline import Pipeline             # route, retrieve, answer, checkclass AskIn(BaseModel):    question: str = Field(min_length=1, max_length=4000)    conversation_id: str | None = None@asynccontextmanagerasync def lifespan(app: FastAPI):    app.state.pipeline = Pipeline.load(settings)   # embedder, BM25, reranker, NLI    yieldapp = FastAPI(lifespan=lifespan)@app.get("/healthz")def healthz() -> dict:    return {"ok": True}                            # the process is alive@app.get("/readyz")def readyz() -> dict:    if not hasattr(app.state, "pipeline"):        raise HTTPException(503, "loading")        # do not send traffic yet    return {"release": settings.release}@app.post("/ask")async def ask(body: AskIn, user: User = Depends(current_user)) -> dict:    pipeline = app.state.pipeline    try:        result = await asyncio.wait_for(            asyncio.to_thread(pipeline.answer, user, body.question, body.conversation_id),            timeout=settings.request_timeout_s)    # 25 s    except asyncio.TimeoutError:        raise HTTPException(504, "PolicyPal took too long. Please try again.")    return result.model_dump() | {"release": settings.release}

Models load once, in the lifespan function, before the service reports itself ready. Loading takes about 40 seconds, which is why there are two health endpoints: /healthz says the process is alive, and /readyz says it can take traffic. The load balancer only sends requests to replicas that are ready.

The user comes from a dependency that validates the company's single sign-on token. Everything downstream, the country filter, the leave tool and the audit log, uses that User, never anything from the request body. The question has a hard length limit, enforced by Pydantic before any model sees it.

pipeline.answer is ordinary blocking code: the reranker and NLI model are CPU-bound, and the model client is synchronous. Running it with asyncio.to_thread keeps the event loop free for other requests. The 25-second timeout gives the user a clear message instead of a hanging page. One caution: the timeout stops waiting, but the thread keeps running until it finishes, so the model client's own 30-second timeout from Section 1 is what finally bounds the work.

Capacity from real numbers

How many requests does PolicyPal handle at once? Little's law gives the answer: the average number of requests in flight equals the arrival rate times the time each one takes. On a normal peak, 30 questions a minute is 0.5 per second, and each takes about 4.5 seconds, so about 2 to 3 requests are in flight. Three replicas are about resilience, not load.

The day that matters is an announcement. When HR emailed all staff about a new hybrid-work policy, questions arrived at 200 a minute for half an hour: about 3.3 per second, or 15 requests in flight. The API replicas coped easily. The limit that bit was the provider's rate limit, counted in tokens per minute: 200 questions a minute at 3,000 input tokens each is 600,000 input tokens a minute, above the account's tier at the time. Requests started receiving 429 errors, retries piled up, and latency climbed. The fix was part capacity (a higher tier, arranged in advance) and part design (the caching in the next lesson).

Degraded modes, decided in advance

Every dependency will fail at some point. Decide what PolicyPal does in each case before it happens, and test each mode.

FailureDegraded behaviourCost to users
Router service downPrompted routerAbout 850 ms slower
Reranker too slowSkip reranking, use hybrid top 5Recall@5 falls from 0.91 to 0.84
Answering model errors or times outRetry once, then a second, smaller modelSlightly lower quality
Provider fully unavailableSearch-only mode: show the top 3 policy sections with linksNo written answer
HRMS downTool returns an error; the model explainsNo balances for a while

Search-only mode deserves a mention. When no model is available, PolicyPal still does something useful: it shows the three best-matching policy sections, reranked, with a note that answers are temporarily unavailable. Employees used it twice in the first quarter and barely noticed.

A release is more than code

A PolicyPal answer depends on six things: the code, the answering model, the prompt, the example bank, the router adapter and the index. Change any one and answers change. So a release pins all six in a manifest, and every answer is logged with the release id.

YAML
# releases/2026.03.18-2.yamlrelease: 2026.03.18-2code: git 4f2c9abanswer_model: claude-sonnet-5fallback_model: claude-haiku-4-5judge_model: claude-opus-5prompt: answer-v2.4examples: bank-v7router_adapter: router-2026-03-02index: policies-2026-03-18T02-00Z     # 6,214 chunks from 397 documents

The nightly index build writes a new, immutable index directory with a timestamped name. It is not used until a release manifest points to it, after the retrieval metrics from Section 6 pass on the new index. Switching releases is a configuration change that replicas pick up on restart; rolling back is pointing at the previous manifest. Nothing is edited in place.

This is the difference between "the answers got worse this week" and "release 2026.03.18-2 changed the example bank from v6 to v7; here are the 11 eval cases it broke".

Check your understanding

0 of 3 answered

1.At a normal peak of 30 questions a minute, each taking 4.5 seconds, about how many requests are in flight at once?

2.Why does PolicyPal have both /healthz and /readyz endpoints?

3.An answer from three weeks ago is reported as wrong. What makes it possible to find the cause quickly?