Applied AI Engineering: From Prompt to Production

Course Content

Applied AI Engineering: From Prompt to Production

9 sections · 29 lessons

Metrics for RAG and agents


One Monday, the eval suite reported that PolicyPal's pass rate had dropped four points over the weekend's changes. Three things had changed: the nightly index rebuild picked up 12 updated policies, someone edited the example bank, and the reranker library was upgraded. Which one did it?

With a single pass-or-fail number per case, finding out meant reading dozens of answers. With metrics for each stage of the pipeline, the answer was on the report in the first minute: retrieval was unchanged, faithfulness was unchanged, but the rate of correct not_in_policy decisions had fallen from 90% to 70%. That pointed straight at the example bank.

PolicyPal is a pipeline. A route is chosen, chunks are retrieved, an answer is generated and checked, and sometimes tools are called in a loop. Each stage can fail on its own, and each needs its own measurement.

130 answerable cases, retrieval crossed with correctness108: as designed10: generation3: lucky,check labels9: retrievalAnswer rightAnswer wrongChunk in top 5Chunk missing
Crossing retrieval success with answer correctness tells you which team owns each failure.

Retrieval metrics

You met recall@k in Section 3: the share of questions where at least one correct chunk is in the top k. It answers "did the answer reach the model at all?". Mean reciprocal rank (MRR) adds position: for each question, score 1 divided by the rank of the first correct chunk, or 0 if none, and average. A correct chunk at rank 1 scores 1.0; at rank 4, 0.25. MRR tells you whether the right chunk is near the top, which matters because models pay more attention to the first sources they read.

Python
# policypal/evals/metrics.pydef first_hit(results: list[dict], gold: list[dict]) -> int | None:    for rank, chunk in enumerate(results, start=1):        section = chunk["section"].split()[0]           # "4.2 Carry forward ..." -> "4.2"        if any(chunk["file"] == g["file"] and section == g["section"] for g in gold):            return rank    return Nonedef retrieval_metrics(runs: list[tuple[list[dict], list[dict]]], k: int = 5) -> dict:    """runs: (ranked results, gold sources) for each case with an answer."""    ranks = [first_hit(results, gold) for results, gold in runs]    return {f"recall@{k}": sum(r is not None and r <= k for r in ranks) / len(ranks),            "mrr": sum(1 / r for r in ranks if r) / len(ranks)}

Only cases that have an answer in the policies count here. For PolicyPal's 130 such cases, the current pipeline scores recall@5 of 0.91 and MRR of 0.78.

Generation metrics

Once the right text is in the prompt, three questions remain.

  • Faithfulness: what share of answer sentences are supported by the retrieved sources? This is the NLI check from Section 3, applied offline. It measures whether the model stayed inside its evidence, regardless of whether the evidence was right.
  • Key-fact correctness: does the answer contain every key fact and none of the must-not statements? This compares the answer to the truth.
  • Status correctness: did the model choose answered, not_in_policy, needs_human or out_of_scope correctly?

Key facts can be checked with the same NLI model, this time with the answer as the premise and each key fact as the hypothesis. It is strict about paraphrase, so PolicyPal also uses the calibrated judge from the next lesson, and counts a fact as present if either says so.

The power of separate metrics comes from combining them. For each case with an answer, cross retrieval success with answer correctness.

Answer correctAnswer wrong
Right chunk in top 5108: working as designed10: a generation problem
Right chunk missing3: lucky, or answered from another section9: a retrieval problem

Each cell has a different owner and a different fix. The 10 cases in the top right are prompt or model work. The 9 in the bottom right are chunking or search work. The 3 lucky cases are worth reading too, since the gold labels may be incomplete.

Agent metrics

An agent run is graded on where it ended and how it got there. The eval harness records every tool call the agent makes, through a thin wrapper around run_tool shown below, and each tool case lists the call it expects, with only the arguments that matter.

Python
# policypal/evals/agent_metrics.pySIDE_EFFECTS = {"create_it_ticket"}def agent_metrics(case: dict, executed: list[dict], pending: dict | None,                  steps: int, tokens: int) -> dict:    expected = case.get("expected_call")          # e.g. {"name": ..., "input": {"urgency": "normal"}}    made = executed + ([pending] if pending else [])    match = [c for c in made if expected and c["name"] == expected["name"]]    return {        "right_tool": bool(match) if expected else not any(c["name"] in SIDE_EFFECTS for c in made),        "right_args": bool(match) and all(match[0]["input"].get(k) == v                                          for k, v in expected["input"].items()),        "extra_calls": max(0, len(made) - (1 if expected else 0)),        "unconfirmed_side_effect": any(c["name"] in SIDE_EFFECTS for c in executed),        "steps": steps, "tokens": tokens,    }

Four kinds of number come out. Outcome: was the task done, with the right tool and arguments? Safety: unconfirmed_side_effect must be zero on every run, every time; one is a release blocker. Efficiency: extra calls, steps and tokens, which turn into latency and cost. And the final answer, graded with the same key-fact checks as any other answer.

Agents need one more habit: run each agent case several times. The same input can take different paths on different runs. PolicyPal runs each of its 60 tool and multi-intent cases three times and reports the share of all runs that succeed, and separately the share of cases that succeed on every run. A case that passes two runs out of three is a flaky feature, not a passing one.

Recording what the agent really did

agent_metrics needs the calls that actually ran, not the calls the final answer mentions. The agent loop from Section 4 does not return that list, and production has no use for it. So the harness wraps run_tool for the length of one case and writes down every call that passes through.

Python
# policypal/evals/recorder.pyfrom contextlib import contextmanagerfrom dataclasses import dataclassfrom unittest import mockfrom policypal import agent, toolsfrom policypal.evals.agent_metrics import agent_metricsLIVE_OK = {"search_policies", "calculate_accrual"}    # read-only: safe to run for real@dataclassclass EvalUser:                         # a test identity; it exists in no live system    employee_id: str    email: str    country: str    grade: str@dataclassclass CaseRun:    text: str    pending: dict | None    executed: list[dict]    steps: int    tokens: int

EvalUser has the attributes the tool handlers read from a session user, so the tools cannot tell the difference. CaseRun holds what every grader needs; Section 9's red-team suite reads pending and executed from the same object.

Python
# policypal/evals/recorder.py (continued)def _returns(payload: dict):    return lambda user, args: payload@contextmanagerdef recording_tools(fixtures: dict[str, dict] | None = None):    """Record every tool call the agent executes; fixtures stand in for live systems."""    executed: list[dict] = []    real_run_tool = agent.run_tool    def recorded(user, call: dict) -> dict:        result = real_run_tool(user, call)             # still validates the input        executed.append({"name": call["name"], "input": dict(call["input"]),                         "is_error": result["is_error"]})        return result    fixtures = fixtures or {}    fakes = {name: (schema, _returns(fixtures.get(name, {"eval_stub": True})))             for name, (schema, _) in tools.HANDLERS.items()             if name in fixtures or name not in LIVE_OK}    with mock.patch.dict(tools.HANDLERS, fakes), mock.patch.object(agent, "run_tool", recorded):        yield executeddef run_agent_case(llm, system: str, case: dict) -> CaseRun:    user = EvalUser(employee_id="eval-0001", email="eval-bot@harbourline.example",                    **case["employee"])    with recording_tools(case.get("fixtures")) as executed:        result = agent.run_agent(llm, system, [{"role": "user", "content": case["input"]}], user)    return CaseRun(result.text, result.pending, executed, result.steps, result.tokens)def score_agent_case(llm, system: str, case: dict, runs: int = 3) -> list[dict]:    return [agent_metrics(case, r.executed, r.pending, r.steps, r.tokens)            for r in (run_agent_case(llm, system, case) for _ in range(runs))]

Four details here are easy to get wrong.

Patch the name the loop uses. agent.py says from policypal.tools import run_tool, which copies the function into the agent module. Patching tools.run_tool would change nothing the loop sees. The recorder would record zero calls, and unconfirmed_side_effect would pass on every run, forever. So the harness also has a self-test: a scripted FakeLLM asks for one search, and the test asserts the recorder saw exactly one call.

Record through the real executor. The wrapper calls the real run_tool, so validation and error results behave as in production, and is_error lets the report count calls the model got wrong.

Nothing live that can change. Only tools on LIVE_OK run for real; search does, because retrieval is part of what is measured. Every other tool returns the case's fixture or a stub. The list names what may run, so a write tool added next quarter is faked by default. If a bug ever lets the loop run create_it_ticket without a click, the eval records it as a side effect and the helpdesk receives no ticket.

Fixtures make balance cases repeatable. The eval user exists in no HRMS, and real balances change every month, so a balance case carries its answer: "fixtures": {"get_leave_balance": {"leave_type": "earned", "available_days": 12, "pending_days": 2, "as_of": "2026-03-14"}}. Only the handler is replaced, so a call for "annual" leave still fails validation.

mock.patch undoes both patches when the block ends, even after an exception. Patching a module is not safe across threads, so each worker process runs its cases one at a time; four processes finish the 180 agent runs in about five minutes.

One report, by stage and group

All these metrics come together in one report per run, compared with the last release.

Metricv2.3v2.4Change
Retrieval recall@50.910.910
Retrieval MRR0.780.79+0.01
Faithfulness (sentences supported)97.2%97.0%-0.2
Key-fact correctness86.9%84.6%-2.3
Status correct, not_in_policy cases90%70%-20
Agent task success (all runs)91.7%92.2%+0.5
Unconfirmed side effects000
Median cost per case$0.0089$0.0090+1%

Reading top to bottom localises the problem: retrieval is fine, faithfulness is fine, one status is not. That row points at whatever teaches the model when to say "not in the policy", which is the example bank.

Check your understanding

0 of 3 answered

1.Faithfulness is 97% but key-fact correctness is only 87%. What does this pattern most likely mean?

2.Why does PolicyPal run each agent case three times?

3.A case has the right chunk in the top 5, but the answer is wrong. Who should look at it first?