Course Content
LangGraph Agents
7 sections · 49 lessons
How can multi-agent graphs be evaluated for task success, tool efficiency, and result quality?
What you need to know
| Level | Metrics | How |
|---|---|---|
| Task success | Success rate, rubric score | 50–300 real cases with expected outcomes |
| Trajectory | Tool calls, redundant calls, handoffs, loops, budget stops | Traces compared with reference trajectories |
| Components | Router accuracy, retrieval recall, critic agreement with humans | Labelled sets per node |
| Economics | Cost and latency per successful task, p95 latency | Token usage from traces |
Tools
- Tracing — LangSmith traces every node, tool call and handoff as a span; OpenTelemetry export works too.
- Trajectory evaluators — the open-source
agentevalspackage compares an agent's tool-call sequence with a reference (strict, unordered, subset or superset match) or uses an LLM judge on the trajectory. - Output evaluators —
openevalsand LangSmith evaluators for correctness, groundedness and rubric scoring. - State history —
get_state_historyon failed threads shows exactly where the run went wrong.
Practices
- Keep a single-agent baseline in every report.
- Freeze the dataset and version it; add every production failure as a new case.
- Run several trials per case; agents are non-deterministic, so report the average and the spread.
- Gate releases in CI: block a change that lowers success or raises cost per success beyond a threshold.
A real-life example
An insurer's claims assistant uses three agents. The eval set is 200 past claims with known outcomes. Release 1.4 changed only the supervisor prompt. Task success stayed at 88%, but trajectory metrics showed average handoffs rising from 2.1 to 3.4 and cost per successful claim from Rs 11 to Rs 17 — the supervisor now asked the document agent to re-read files it had already processed. Because the CI gate checked cost per success, the change was blocked before release, and the prompt was fixed in a day.
Follow-up questions to expect
- "How do you evaluate without expected answers?" — Rubrics scored by an LLM judge that is checked against human labels on a sample, plus pairwise comparison between versions.
- "How many cases are enough?" — Enough to see the differences you care about; 100–300 well-chosen cases catch most regressions, with more for rare, risky categories.
- "Offline or online evaluation?" — Both. Offline suites gate releases; online monitoring on sampled production traces catches drift and new failure types.