LangGraph Agents

Course Content

LangGraph Agents

7 sections · 49 lessons

How can multi-agent graphs be evaluated for task success, tool efficiency, and result quality?


Release 1.4 changed only the supervisor prompt88%88%2.13.4Rs 11Rs 171.31.4task successhandoffs per claimcost per success
Answer-only evals would have shipped this release; trajectory and cost-per-success metrics blocked it.

What you need to know

LevelMetricsHow
Task successSuccess rate, rubric score50–300 real cases with expected outcomes
TrajectoryTool calls, redundant calls, handoffs, loops, budget stopsTraces compared with reference trajectories
ComponentsRouter accuracy, retrieval recall, critic agreement with humansLabelled sets per node
EconomicsCost and latency per successful task, p95 latencyToken usage from traces

Tools

  • Tracing — LangSmith traces every node, tool call and handoff as a span; OpenTelemetry export works too.
  • Trajectory evaluators — the open-source agentevals package compares an agent's tool-call sequence with a reference (strict, unordered, subset or superset match) or uses an LLM judge on the trajectory.
  • Output evaluators — openevals and LangSmith evaluators for correctness, groundedness and rubric scoring.
  • State history — get_state_history on failed threads shows exactly where the run went wrong.

Practices

  • Keep a single-agent baseline in every report.
  • Freeze the dataset and version it; add every production failure as a new case.
  • Run several trials per case; agents are non-deterministic, so report the average and the spread.
  • Gate releases in CI: block a change that lowers success or raises cost per success beyond a threshold.

A real-life example

An insurer's claims assistant uses three agents. The eval set is 200 past claims with known outcomes. Release 1.4 changed only the supervisor prompt. Task success stayed at 88%, but trajectory metrics showed average handoffs rising from 2.1 to 3.4 and cost per successful claim from Rs 11 to Rs 17 — the supervisor now asked the document agent to re-read files it had already processed. Because the CI gate checked cost per success, the change was blocked before release, and the prompt was fixed in a day.

Follow-up questions to expect

  • "How do you evaluate without expected answers?" — Rubrics scored by an LLM judge that is checked against human labels on a sample, plus pairwise comparison between versions.
  • "How many cases are enough?" — Enough to see the differences you care about; 100–300 well-chosen cases catch most regressions, with more for rare, risky categories.
  • "Offline or online evaluation?" — Both. Offline suites gate releases; online monitoring on sampled production traces catches drift and new failure types.