Course Content
Agentic AI Patterns
9 sections · 50 lessons
How do you evaluate the performance of an AI agent system?
What you need to know
The three levels
| Level | Question | Example checks |
|---|---|---|
| Component | Did each part work? | Retrieval recall at 5; tool-choice accuracy; argument validity |
| Trajectory | Was the path sensible and safe? | Required tools called; forbidden tools not called; no repeated calls; steps within budget |
| Outcome | Was the task done? | Final database state correct; tests pass; ticket resolved; cost and latency |
Why trajectory matters
Two runs can give the same correct answer, where one checked the policy and the other guessed. The second is a time bomb. Trajectory checks catch "right answer, wrong process", and also safety problems like calling approve_claim before get_prior_claims.
1def score_trajectory(calls, must_call, must_not_call, max_steps):2 names = [c["name"] for c in calls]3 repeats = sum(1 for a, b in zip(calls, calls[1:])4 if a["name"] == b["name"] and a["args"] == b["args"])5 return {6 "required_tools_used": all(t in names for t in must_call),7 "forbidden_tool_used": any(t in names for t in must_not_call),8 "steps": len(calls),9 "within_budget": len(calls) <= max_steps,10 "identical_repeats": repeats, # same call twice in a row = likely loop11 }Each golden case lists the tools that must be called and the ones that must not. A run that repeats get_policy and calls approve_claim on a case that should be escalated scores identical_repeats: 1 and forbidden_tool_used: True.
Method
- Golden set from real traces, over-sampling failures and edge cases.
- Run in CI on every prompt, model or tool change.
- Code checks first: final state, schema, arithmetic, required steps.
- LLM judges only where no code check exists, with a written rubric, calibrated against human labels, and their agreement rate reported.
- Several runs per case. Agents are non-deterministic. A useful stricter metric, pass^k, asks whether the agent succeeds on all k tries of the same task, which measures consistency, not just luck.
- Online: A/B tests, escalation and override rates, reopen rates, user feedback.
A real-life example
A travel-booking agent has 150 golden cases. Each case has a final-state check (the booking in the sandbox has the right dates, fare class and traveller), required tools (get_travel_policy before any hold_booking), and forbidden tools for some cases (no confirm_booking when the fare exceeds the cap without approval).
Version A: 88% outcome success, $0.09 per task, median 7 steps. Version B, with a stronger model: 91% success, $0.31 per task, median 6 steps. Running each case 4 times showed version A succeeded on all 4 tries in 71% of cases, and B in 84%.
The team shipped a mix: B only for multi-city trips, where the consistency gain mattered, and A for simple return trips. Cost per successful task fell compared with using B everywhere.
Follow-up questions to expect
- "How do you trust an LLM judge?" — Have humans label 100 or so outputs, measure the judge's agreement with them, fix the rubric until agreement is high, and re-check after model changes.
- "What if there are many valid paths?" — Do not require an exact sequence. Check required and forbidden steps, ordering constraints that matter, and the final state.
- "What is the single metric for stakeholders?" — Cost per successful task, alongside success rate. It captures quality and economics together.