Course Content
LLM Evaluation
6 sections · 50 lessons
How can you evaluate the performance of AI agents beyond just final outputs?
What you need to know
An agent can reach the right answer by luck after ten wasted calls, or reach the wrong answer because one tool argument was wrong in step 2. Scoring only the final text hides both.
What to measure
| Metric | Question it answers |
|---|---|
| Task success | Is the world in the goal state? (row inserted, ticket closed, tests pass) |
| Tool selection | Did it pick the right tool at each step? |
| Argument correctness | Were arguments valid and grounded in the request, not invented? |
| Step efficiency | Steps and tokens used vs the minimum needed |
| Recovery | After a tool error, did it adapt or repeat the same failing call? |
| Safety and scope | Any action it was not allowed or asked to take? |
| Cost and latency per completed task | Not per call — a cheap agent that fails half the time is expensive |
Reliability across repeats
Agents are non-deterministic, so run each task several times. The τ-bench benchmark popularised pass^k: the chance that the agent succeeds on all k tries of the same task. If one try succeeds 80% of the time and tries are independent, pass^4 = 0.8 to the power 4 ≈ 0.41. An agent that looks good on average can be unreliable for any single customer.
How to grade the path
- Outcome checks — query the sandbox database or run the test suite. Most trustworthy.
- Reference path matching — compare tool calls to an expected list; use loose matching (right tools in any order, or a subset) because many valid paths exist.
- Per-step judge — ask "given the state so far, was this step reasonable?" to find the first bad step.
Tools that help: LangSmith trajectory evaluators, DeepEval's tool-correctness and task-completion metrics, and Inspect, which runs agent tasks in sandboxed containers with scorers.
A real-life example
The code-review bot is now an agent with four tools: get_diff, search_repo, run_tests and post_comment. On a task where a PR removes a null check, a good trajectory is: get the diff, search for callers of the function, run the tests, post one comment on the right line.
Across 50 seeded PRs, run 3 times each, final-answer accuracy is 78%. The trajectory view shows more: in 11% of runs the agent called run_tests before reading the diff and then commented on unrelated failing tests; in 6% it looped on search_repo with the same query until it hit the step limit. Fixing the tool description for search_repo cut the loops to under 1% — a fix you cannot find from the final comment alone.
Follow-up questions to expect
- "Why not match the exact reference trajectory?" — Many plans are valid; exact matching fails good agents that took a different route. Verify the outcome and check for forbidden or wasteful steps instead.
- "How do you make agent evals repeatable?" — A sandbox with seeded state, mocked or recorded external APIs, pinned model versions, and full traces logged for replay.
- "What is pass^k vs pass@k?" — pass@k is the chance at least one of k tries succeeds (useful with a verifier); pass^k is the chance all k succeed (reliability for users).