LLM Evaluation

Course Content

LLM Evaluation

6 sections · 50 lessons

How can you evaluate the performance of AI agents beyond just final outputs?


One code-review trajectory, graded step by stepget_diffsearch_repo:callersrun_testspost_commenton line 42nulllooped herein 6% of runscalled first in 11%Final comments were right 78% of the time; the path showed where the other runs went wrong.
Grading only the last step hides the loop and the wrong order that a trajectory makes obvious.

What you need to know

An agent can reach the right answer by luck after ten wasted calls, or reach the wrong answer because one tool argument was wrong in step 2. Scoring only the final text hides both.

What to measure

MetricQuestion it answers
Task successIs the world in the goal state? (row inserted, ticket closed, tests pass)
Tool selectionDid it pick the right tool at each step?
Argument correctnessWere arguments valid and grounded in the request, not invented?
Step efficiencySteps and tokens used vs the minimum needed
RecoveryAfter a tool error, did it adapt or repeat the same failing call?
Safety and scopeAny action it was not allowed or asked to take?
Cost and latency per completed taskNot per call — a cheap agent that fails half the time is expensive

Reliability across repeats

Agents are non-deterministic, so run each task several times. The τ-bench benchmark popularised pass^k: the chance that the agent succeeds on all k tries of the same task. If one try succeeds 80% of the time and tries are independent, pass^4 = 0.8 to the power 4 ≈ 0.41. An agent that looks good on average can be unreliable for any single customer.

How to grade the path

  • Outcome checks — query the sandbox database or run the test suite. Most trustworthy.
  • Reference path matching — compare tool calls to an expected list; use loose matching (right tools in any order, or a subset) because many valid paths exist.
  • Per-step judge — ask "given the state so far, was this step reasonable?" to find the first bad step.

Tools that help: LangSmith trajectory evaluators, DeepEval's tool-correctness and task-completion metrics, and Inspect, which runs agent tasks in sandboxed containers with scorers.

A real-life example

The code-review bot is now an agent with four tools: get_diff, search_repo, run_tests and post_comment. On a task where a PR removes a null check, a good trajectory is: get the diff, search for callers of the function, run the tests, post one comment on the right line.

Across 50 seeded PRs, run 3 times each, final-answer accuracy is 78%. The trajectory view shows more: in 11% of runs the agent called run_tests before reading the diff and then commented on unrelated failing tests; in 6% it looped on search_repo with the same query until it hit the step limit. Fixing the tool description for search_repo cut the loops to under 1% — a fix you cannot find from the final comment alone.

Follow-up questions to expect

  • "Why not match the exact reference trajectory?" — Many plans are valid; exact matching fails good agents that took a different route. Verify the outcome and check for forbidden or wasteful steps instead.
  • "How do you make agent evals repeatable?" — A sandbox with seeded state, mocked or recorded external APIs, pinned model versions, and full traces logged for replay.
  • "What is pass^k vs pass@k?" — pass@k is the chance at least one of k tries succeeds (useful with a verifier); pass^k is the chance all k succeed (reliability for users).