Agentic AI Patterns

Course Content

Agentic AI Patterns

9 sections · 50 lessons

How do you evaluate the performance of an AI agent system?


Three levels, three different fixesComponent — right tool, valid argsTrajectory — required steps, no loopsOutcome — final state, cost, latency
A correct answer reached by skipping the policy check passes the outcome test and fails the trajectory one — and that is the run that hurts later.

What you need to know

The three levels

LevelQuestionExample checks
ComponentDid each part work?Retrieval recall at 5; tool-choice accuracy; argument validity
TrajectoryWas the path sensible and safe?Required tools called; forbidden tools not called; no repeated calls; steps within budget
OutcomeWas the task done?Final database state correct; tests pass; ticket resolved; cost and latency

Why trajectory matters

Two runs can give the same correct answer, where one checked the policy and the other guessed. The second is a time bomb. Trajectory checks catch "right answer, wrong process", and also safety problems like calling approve_claim before get_prior_claims.

Python
def score_trajectory(calls, must_call, must_not_call, max_steps):    names = [c["name"] for c in calls]    repeats = sum(1 for a, b in zip(calls, calls[1:])                  if a["name"] == b["name"] and a["args"] == b["args"])    return {        "required_tools_used": all(t in names for t in must_call),        "forbidden_tool_used": any(t in names for t in must_not_call),        "steps": len(calls),        "within_budget": len(calls) <= max_steps,        "identical_repeats": repeats,       # same call twice in a row = likely loop    }

Each golden case lists the tools that must be called and the ones that must not. A run that repeats get_policy and calls approve_claim on a case that should be escalated scores identical_repeats: 1 and forbidden_tool_used: True.

Method

  • Golden set from real traces, over-sampling failures and edge cases.
  • Run in CI on every prompt, model or tool change.
  • Code checks first: final state, schema, arithmetic, required steps.
  • LLM judges only where no code check exists, with a written rubric, calibrated against human labels, and their agreement rate reported.
  • Several runs per case. Agents are non-deterministic. A useful stricter metric, pass^k, asks whether the agent succeeds on all k tries of the same task, which measures consistency, not just luck.
  • Online: A/B tests, escalation and override rates, reopen rates, user feedback.

A real-life example

A travel-booking agent has 150 golden cases. Each case has a final-state check (the booking in the sandbox has the right dates, fare class and traveller), required tools (get_travel_policy before any hold_booking), and forbidden tools for some cases (no confirm_booking when the fare exceeds the cap without approval).

Version A: 88% outcome success, $0.09 per task, median 7 steps. Version B, with a stronger model: 91% success, $0.31 per task, median 6 steps. Running each case 4 times showed version A succeeded on all 4 tries in 71% of cases, and B in 84%.

The team shipped a mix: B only for multi-city trips, where the consistency gain mattered, and A for simple return trips. Cost per successful task fell compared with using B everywhere.

Follow-up questions to expect

  • "How do you trust an LLM judge?" — Have humans label 100 or so outputs, measure the judge's agreement with them, fix the rubric until agreement is high, and re-check after model changes.
  • "What if there are many valid paths?" — Do not require an exact sequence. Check required and forbidden steps, ordering constraints that matter, and the final state.
  • "What is the single metric for stakeholders?" — Cost per successful task, alongside success rate. It captures quality and economics together.