Course Content
Applied AI Engineering: From Prompt to Production
9 sections · 29 lessons
Metrics for RAG and agents
One Monday, the eval suite reported that PolicyPal's pass rate had dropped four points over the weekend's changes. Three things had changed: the nightly index rebuild picked up 12 updated policies, someone edited the example bank, and the reranker library was upgraded. Which one did it?
With a single pass-or-fail number per case, finding out meant reading dozens of answers. With metrics for each stage of the pipeline, the answer was on the report in the first minute: retrieval was unchanged, faithfulness was unchanged, but the rate of correct not_in_policy decisions had fallen from 90% to 70%. That pointed straight at the example bank.
PolicyPal is a pipeline. A route is chosen, chunks are retrieved, an answer is generated and checked, and sometimes tools are called in a loop. Each stage can fail on its own, and each needs its own measurement.
Retrieval metrics
You met recall@k in Section 3: the share of questions where at least one correct chunk is in the top k. It answers "did the answer reach the model at all?". Mean reciprocal rank (MRR) adds position: for each question, score 1 divided by the rank of the first correct chunk, or 0 if none, and average. A correct chunk at rank 1 scores 1.0; at rank 4, 0.25. MRR tells you whether the right chunk is near the top, which matters because models pay more attention to the first sources they read.
1# policypal/evals/metrics.py2def first_hit(results: list[dict], gold: list[dict]) -> int | None:3 for rank, chunk in enumerate(results, start=1):4 section = chunk["section"].split()[0] # "4.2 Carry forward ..." -> "4.2"5 if any(chunk["file"] == g["file"] and section == g["section"] for g in gold):6 return rank7 return None89def retrieval_metrics(runs: list[tuple[list[dict], list[dict]]], k: int = 5) -> dict:10 """runs: (ranked results, gold sources) for each case with an answer."""11 ranks = [first_hit(results, gold) for results, gold in runs]12 return {f"recall@{k}": sum(r is not None and r <= k for r in ranks) / len(ranks),13 "mrr": sum(1 / r for r in ranks if r) / len(ranks)}Only cases that have an answer in the policies count here. For PolicyPal's 130 such cases, the current pipeline scores recall@5 of 0.91 and MRR of 0.78.
Generation metrics
Once the right text is in the prompt, three questions remain.
- Faithfulness: what share of answer sentences are supported by the retrieved sources? This is the NLI check from Section 3, applied offline. It measures whether the model stayed inside its evidence, regardless of whether the evidence was right.
- Key-fact correctness: does the answer contain every key fact and none of the must-not statements? This compares the answer to the truth.
- Status correctness: did the model choose
answered,not_in_policy,needs_humanorout_of_scopecorrectly?
Key facts can be checked with the same NLI model, this time with the answer as the premise and each key fact as the hypothesis. It is strict about paraphrase, so PolicyPal also uses the calibrated judge from the next lesson, and counts a fact as present if either says so.
The power of separate metrics comes from combining them. For each case with an answer, cross retrieval success with answer correctness.
| Answer correct | Answer wrong | |
|---|---|---|
| Right chunk in top 5 | 108: working as designed | 10: a generation problem |
| Right chunk missing | 3: lucky, or answered from another section | 9: a retrieval problem |
Each cell has a different owner and a different fix. The 10 cases in the top right are prompt or model work. The 9 in the bottom right are chunking or search work. The 3 lucky cases are worth reading too, since the gold labels may be incomplete.
Agent metrics
An agent run is graded on where it ended and how it got there. The eval harness records every tool call the agent makes, through a thin wrapper around run_tool shown below, and each tool case lists the call it expects, with only the arguments that matter.
1# policypal/evals/agent_metrics.py2SIDE_EFFECTS = {"create_it_ticket"}34def agent_metrics(case: dict, executed: list[dict], pending: dict | None,5 steps: int, tokens: int) -> dict:6 expected = case.get("expected_call") # e.g. {"name": ..., "input": {"urgency": "normal"}}7 made = executed + ([pending] if pending else [])8 match = [c for c in made if expected and c["name"] == expected["name"]]9 return {10 "right_tool": bool(match) if expected else not any(c["name"] in SIDE_EFFECTS for c in made),11 "right_args": bool(match) and all(match[0]["input"].get(k) == v12 for k, v in expected["input"].items()),13 "extra_calls": max(0, len(made) - (1 if expected else 0)),14 "unconfirmed_side_effect": any(c["name"] in SIDE_EFFECTS for c in executed),15 "steps": steps, "tokens": tokens,16 }Four kinds of number come out. Outcome: was the task done, with the right tool and arguments? Safety: unconfirmed_side_effect must be zero on every run, every time; one is a release blocker. Efficiency: extra calls, steps and tokens, which turn into latency and cost. And the final answer, graded with the same key-fact checks as any other answer.
Agents need one more habit: run each agent case several times. The same input can take different paths on different runs. PolicyPal runs each of its 60 tool and multi-intent cases three times and reports the share of all runs that succeed, and separately the share of cases that succeed on every run. A case that passes two runs out of three is a flaky feature, not a passing one.
Recording what the agent really did
agent_metrics needs the calls that actually ran, not the calls the final answer mentions. The agent loop from Section 4 does not return that list, and production has no use for it. So the harness wraps run_tool for the length of one case and writes down every call that passes through.
1# policypal/evals/recorder.py2from contextlib import contextmanager3from dataclasses import dataclass4from unittest import mock56from policypal import agent, tools7from policypal.evals.agent_metrics import agent_metrics89LIVE_OK = {"search_policies", "calculate_accrual"} # read-only: safe to run for real1011@dataclass12class EvalUser: # a test identity; it exists in no live system13 employee_id: str14 email: str15 country: str16 grade: str1718@dataclass19class CaseRun:20 text: str21 pending: dict | None22 executed: list[dict]23 steps: int24 tokens: intEvalUser has the attributes the tool handlers read from a session user, so the tools cannot tell the difference. CaseRun holds what every grader needs; Section 9's red-team suite reads pending and executed from the same object.
1# policypal/evals/recorder.py (continued)2def _returns(payload: dict):3 return lambda user, args: payload45@contextmanager6def recording_tools(fixtures: dict[str, dict] | None = None):7 """Record every tool call the agent executes; fixtures stand in for live systems."""8 executed: list[dict] = []9 real_run_tool = agent.run_tool1011 def recorded(user, call: dict) -> dict:12 result = real_run_tool(user, call) # still validates the input13 executed.append({"name": call["name"], "input": dict(call["input"]),14 "is_error": result["is_error"]})15 return result1617 fixtures = fixtures or {}18 fakes = {name: (schema, _returns(fixtures.get(name, {"eval_stub": True})))19 for name, (schema, _) in tools.HANDLERS.items()20 if name in fixtures or name not in LIVE_OK}21 with mock.patch.dict(tools.HANDLERS, fakes), mock.patch.object(agent, "run_tool", recorded):22 yield executed2324def run_agent_case(llm, system: str, case: dict) -> CaseRun:25 user = EvalUser(employee_id="eval-0001", email="eval-bot@harbourline.example",26 **case["employee"])27 with recording_tools(case.get("fixtures")) as executed:28 result = agent.run_agent(llm, system, [{"role": "user", "content": case["input"]}], user)29 return CaseRun(result.text, result.pending, executed, result.steps, result.tokens)3031def score_agent_case(llm, system: str, case: dict, runs: int = 3) -> list[dict]:32 return [agent_metrics(case, r.executed, r.pending, r.steps, r.tokens)33 for r in (run_agent_case(llm, system, case) for _ in range(runs))]Four details here are easy to get wrong.
Patch the name the loop uses. agent.py says from policypal.tools import run_tool, which copies the function into the agent module. Patching tools.run_tool would change nothing the loop sees. The recorder would record zero calls, and unconfirmed_side_effect would pass on every run, forever. So the harness also has a self-test: a scripted FakeLLM asks for one search, and the test asserts the recorder saw exactly one call.
Record through the real executor. The wrapper calls the real run_tool, so validation and error results behave as in production, and is_error lets the report count calls the model got wrong.
Nothing live that can change. Only tools on LIVE_OK run for real; search does, because retrieval is part of what is measured. Every other tool returns the case's fixture or a stub. The list names what may run, so a write tool added next quarter is faked by default. If a bug ever lets the loop run create_it_ticket without a click, the eval records it as a side effect and the helpdesk receives no ticket.
Fixtures make balance cases repeatable. The eval user exists in no HRMS, and real balances change every month, so a balance case carries its answer: "fixtures": {"get_leave_balance": {"leave_type": "earned", "available_days": 12, "pending_days": 2, "as_of": "2026-03-14"}}. Only the handler is replaced, so a call for "annual" leave still fails validation.
mock.patch undoes both patches when the block ends, even after an exception. Patching a module is not safe across threads, so each worker process runs its cases one at a time; four processes finish the 180 agent runs in about five minutes.
One report, by stage and group
All these metrics come together in one report per run, compared with the last release.
| Metric | v2.3 | v2.4 | Change |
|---|---|---|---|
| Retrieval recall@5 | 0.91 | 0.91 | 0 |
| Retrieval MRR | 0.78 | 0.79 | +0.01 |
| Faithfulness (sentences supported) | 97.2% | 97.0% | -0.2 |
| Key-fact correctness | 86.9% | 84.6% | -2.3 |
Status correct, not_in_policy cases | 90% | 70% | -20 |
| Agent task success (all runs) | 91.7% | 92.2% | +0.5 |
| Unconfirmed side effects | 0 | 0 | 0 |
| Median cost per case | $0.0089 | $0.0090 | +1% |
Reading top to bottom localises the problem: retrieval is fine, faithfulness is fine, one status is not. That row points at whatever teaches the model when to say "not in the policy", which is the example bank.
Check your understanding
0 of 3 answered
1.Faithfulness is 97% but key-fact correctness is only 87%. What does this pattern most likely mean?
2.Why does PolicyPal run each agent case three times?
3.A case has the right chunk in the top 5, but the answer is wrong. Who should look at it first?