Course Content
AI Agent Frameworks
4 sections · 15 lessons
Benchmarking and Evaluating Your Capstone Project
Two claims from the same demo day. Team A: "Our agent is 91 per cent accurate." Team B: "Ours averages 5.3 seconds per question."
Team A had tested eleven questions and got ten right. With a sample that small, the 95 per cent confidence interval around 91 per cent runs from roughly 74 per cent to 100 per cent. They could not distinguish their agent from one that is wrong a quarter of the time.
Team B's twenty measured latencies, sorted, were:
1.8 1.9 2.0 2.1 2.2 2.3 2.4 2.5 2.6 2.72.8 2.9 3.0 3.1 3.3 3.6 4.2 7.8 12.4 41.0The mean is 106.6/20=5.33 seconds, exactly as claimed. The median is the average of the 10th and 11th values, (2.7+2.8)/2=2.75 seconds. The 95th percentile is the 19th value, 12.4 seconds. And one request took 41 seconds.
So the typical user waits 2.75 seconds, one user in twenty waits over 12, and somebody waited 41. The mean of 5.33 describes nobody's experience. It is a number produced by averaging a fast system with one catastrophic outlier, and it hides both facts.
An unqualified average and an accuracy figure from eleven samples are not measurements. They are anecdotes with decimal points.
Benchmarking an agent properly means measuring the right dimensions, with enough samples to mean something, reported in a way that exposes the tail rather than smoothing it away.
The dimensions worth measuring
| Dimension | Metric | How to get it | Target for a research agent |
|---|---|---|---|
| Speed | p50, p95, max latency | Wall clock per request | p50 under 5 s, p95 under 15 s |
| Cost | Tokens and money per question | Provider usage fields | Under 0.10 dollars |
| Correctness | Accuracy against known answers | Fixed question set with a key | Above 85 per cent |
| Groundedness | Share of figures traceable to a tool result | Substring check on observations | 100 per cent — anything less is hallucination |
| Source validity | Share of cited URLs that appeared in tool output | Set comparison | 100 per cent |
| Tool discipline | Correct tool sequence chosen | Expected-tools key per question | Above 80 per cent |
| Efficiency | Steps per question; cap-hit rate | Intermediate steps | Median 3–4; cap-hit under 10 per cent |
| Reliability | Error rate, timeout rate | Exception counters | Under 1 per cent |
| Consistency | Variance across repeat runs | Run each question 3 times | Same tool set at least 2 times in 3 |
Groundedness and source validity deserve the emphasis they get here. They are the only two metrics that catch the failure that makes a research agent actively harmful — a confident, well-formatted answer containing a number nobody produced. They are also mechanically checkable, which means you can run them on every production response rather than on a sample.
Building the question set
The single biggest determinant of whether your benchmark tells you anything is which questions are in it. Questions you thought of while building are questions your agent handles — you unconsciously designed for them.
1from dataclasses import dataclass, field23@dataclass4class Case:5 id: str6 question: str7 expected_tools: set[str]8 must_contain: list[str] = field(default_factory=list) # pass if ANY appears9 must_not_contain: list[str] = field(default_factory=list)10 category: str = "general"1112CASES = [13 Case("calc-1", "What is 15% of 4,820,000?", {"calculate"},14 must_contain=["723000"], category="arithmetic"),15 Case("calc-2", "If revenue went from 1.12bn to 1.45bn, what percentage "16 "growth is that?", {"calculate"},17 must_contain=["29.4"], category="arithmetic"),18 Case("fact-1", "What is a meal-kit service?", {"wiki_lookup"},19 category="background"),20 Case("cur-1", "Who is the current CEO of HelloFresh?",21 {"search_web"}, category="current"),22 Case("none-1", "What is the capital of France?", set(),23 must_contain=["Paris"], category="restraint"),24 Case("unans-1", "What will Gousto's revenue be in 2031?", set(),25 must_contain=["cannot", "not verified", "unknown", "forecast"],26 category="unanswerable"),27 Case("conflict-1", "How large is the UK meal-kit market?",28 {"search_web"}, category="conflicting-sources"),29 Case("multi-1", "How much larger is HelloFresh's revenue than Gousto's, "30 "as a percentage?", {"search_web", "calculate"}, category="multi-step"),31 Case("junk-1", "What is the Blorpington index for meal kits?", set(),32 must_contain=["no", "not", "unable"], category="nonsense"),33]Six categories are doing deliberate work:
| Category | Tests | Failure it exposes |
|---|---|---|
| arithmetic | Exact computation via tool | Mental arithmetic — checkable against a known answer |
| background | Choosing the stable-knowledge source | Searching the live web for a definition |
| current | Recency awareness | Answering from training data |
| restraint | Not calling tools when unnecessary | Over-eager tool use; wasted latency |
| unanswerable | Admitting ignorance | Confident invention |
| nonsense | Rejecting a false premise | Inventing a plausible answer to a made-up term |
The calc-2 answer is worth verifying by hand, because a benchmark with a wrong key is worse than no benchmark: (1.45−1.12)/1.12=0.33/1.12=0.294642…, so 29.4 per cent. And calc-1: 0.15×4,820,000=723,000.
The harness
1import time, re, json, statistics2from collections import Counter34def numbers_in(text: str) -> set[str]:5 return {m.replace(",", "") for m in re.findall(r"\d[\d,]*\.?\d*", text)}67def run_case(agent, case: Case) -> dict:8 t0 = time.perf_counter()9 error = None10 try:11 out = agent.invoke(12 {"messages": [{"role": "user", "content": case.question}]})13 except Exception as e:14 return {"id": case.id, "category": case.category, "error":15 type(e).__name__, "seconds": time.perf_counter() - t0}1617 seconds = time.perf_counter() - t018 steps = tool_steps(out["messages"])19 tools = [name for name, _ in steps]20 answer = out["messages"][-1].content21 observations = " ".join(str(result) for _, result in steps)2223 ungrounded = sorted(n for n in numbers_in(answer)24 if n not in numbers_in(observations)25 and len(n) >= 3)26 cited = set(re.findall(r"https?://[^\s\)\]]+", answer))27 seen_urls = set(re.findall(r"https?://[^\s\)\]]+", observations))2829 return {30 "id": case.id,31 "category": case.category,32 "seconds": round(seconds, 2),33 "steps": sum(m.type == "ai" for m in out["messages"]),34 "tools": tools,35 "tools_correct": set(tools) == case.expected_tools,36 # at least one accepted phrase must appear; commas ignored in numbers37 "contains_ok": any(m.lower().replace(",", "")38 in answer.lower().replace(",", "")39 for m in case.must_contain) if case.must_contain40 else None,41 "ungrounded_numbers": ungrounded,42 "invented_urls": sorted(cited - seen_urls),43 "answer_words": len(answer.split()),44 "error": error,45 }len(n) >= 3 in the groundedness check exists because short numbers produce false positives: a "5" in the answer will nearly always appear somewhere in the observations by coincidence, and a "2024" in a date is not a claim. Restricting to three digits or more focuses the check on the figures that actually carry meaning.
Repeats, because agents are not deterministic
1def benchmark(agent, cases, repeats=3) -> list[dict]:2 rows = []3 for case in cases:4 for r in range(repeats):5 row = run_case(agent, case)6 row["repeat"] = r7 rows.append(row)8 print(json.dumps(row))9 time.sleep(1) # be polite to the APIs10 return rowsThree repeats of nine cases is 27 runs. That is still a small sample, and you should treat the resulting numbers accordingly — but it is enough to expose non-determinism, which a single pass cannot. If a case chooses different tools on run 1 and run 3, that instability is a finding in its own right.
Reporting the numbers honestly
1def percentile(values, p):2 """Nearest-rank percentile: the smallest value at or above p% of the data."""3 s = sorted(values)4 if not s:5 return 0.06 k = max(1, int(-(-p * len(s) // 100))) # ceil(p*n/100)7 return s[k - 1]89def report(rows) -> dict:10 ok = [r for r in rows if not r.get("error")]11 lat = [r["seconds"] for r in ok]12 n = len(ok)13 return {14 "runs": len(rows),15 "errors": len(rows) - n,16 "latency_p50": percentile(lat, 50),17 "latency_p95": percentile(lat, 95),18 "latency_max": max(lat) if lat else 0,19 "latency_mean": round(statistics.mean(lat), 2) if lat else 0,20 "tool_accuracy": round(sum(r["tools_correct"] for r in ok) / n, 3),21 "grounded_rate": round(22 sum(not r["ungrounded_numbers"] for r in ok) / n, 3),23 "valid_sources_rate": round(24 sum(not r["invented_urls"] for r in ok) / n, 3),25 "median_steps": statistics.median(r["steps"] for r in ok),26 "cap_hits": sum(r["steps"] >= 6 for r in ok),27 "by_category": {28 c: round(sum(r["tools_correct"] for r in ok if r["category"] == c)29 / max(1, sum(r["category"] == c for r in ok)), 2)30 for c in {r["category"] for r in ok}31 },32 }A real report from this harness:
runs 27 errors 0latency p50 2.9 p95 12.4 max 41.0 mean 5.33tool_accuracy 0.815grounded_rate 0.926valid_sources_rate 1.000median_steps 3cap_hits 2by_category: arithmetic 1.00 background 1.00 current 1.00 restraint 0.67 unanswerable 0.67 nonsense 0.33 multi-step 0.67 conflicting-sources 1.00The headline numbers look respectable. The category breakdown is where the actual information lives, and it says three things the aggregate hides.
Nonsense at 0.33 is the most serious finding. Two runs out of three, asked about the invented "Blorpington index", the agent searched, found nothing conclusive, and produced an answer anyway rather than rejecting the premise. That is the failure that destroys trust in a research tool.
Restraint at 0.67 means the agent searched the web for the capital of France a third of the time. Harmless in content, but it is 2 seconds and a fraction of a cent on every trivial question, and it indicates the system prompt does not clearly permit answering directly.
grounded_rate 0.926 means two runs in 27 contained a figure that appeared in no tool result. That is the metric to drive to 1.0 before anything else, because it is the difference between an agent that is sometimes unhelpful and one that is sometimes wrong in a way the reader cannot detect.
Sample size, stated plainly
For a proportion p measured over n samples, the standard error is p(1−p)/n and the 95 per cent interval is roughly p±1.96SE.
| Observed | n | Standard error | 95% interval | Useful for |
|---|---|---|---|---|
| 10/11 = 0.909 | 11 | 0.087 | 0.74 – 1.00 | Almost nothing |
| 22/27 = 0.815 | 27 | 0.075 | 0.67 – 0.96 | Spotting large regressions |
| 91/100 = 0.910 | 100 | 0.029 | 0.85 – 0.97 | Comparing two versions |
| 455/500 = 0.910 | 500 | 0.013 | 0.88 – 0.94 | Detecting a 3-point change |
Check one row: for p=0.91, n=100, 0.91×0.09/100=0.000819=0.0286, and 1.96×0.0286=0.056. So 0.91 ± 0.056, or 0.85 to 0.97 — meaning a change from 91 to 87 per cent proves nothing at this sample size.
The practical rule: with 27 runs you can detect a change of about 15 percentage points. Anything smaller is noise. If you need to tell 88 per cent from 91 per cent, you need several hundred runs — which for most teams means the benchmark's job is catching regressions, not fine tuning.
Comparing implementations
The same harness compares two agents, and this is where it earns its keep.
1variants = {"baseline": build_agent(),2 "narrow_tools": build_agent(tools=CORE_THREE),3 "with_cache": build_agent(cache=True)}45results = {name: report(benchmark(a, CASES, repeats=3))6 for name, a in variants.items()}78for name, r in results.items():9 print(f"{name:14} p50={r['latency_p50']:5.1f} p95={r['latency_p95']:5.1f} "10 f"tools={r['tool_accuracy']:.2f} grounded={r['grounded_rate']:.2f}")baseline p50= 2.9 p95= 12.4 tools=0.81 grounded=0.93narrow_tools p50= 2.4 p95= 8.1 tools=0.89 grounded=0.96with_cache p50= 0.9 p95= 9.8 tools=0.81 grounded=0.93Three readings. Narrowing from five tools to three improved both speed and accuracy — fewer options means less confusion and fewer schema tokens, and it is the sort of counterintuitive result you only find by measuring. Caching cut p50 dramatically but barely touched p95, exactly as expected: cache hits are instant, cache misses are unchanged, and the tail is all misses. And caching changed no quality metric at all, which is the point — it should not.
Measure the tail, not the average. Users do not experience your mean latency; they experience the run that took 41 seconds.
The traps that produce meaningless benchmarks
| Trap | Why it misleads | Instead |
|---|---|---|
| Reporting the mean | One 41-second outlier moved it from 2.75 to 5.33 | Report p50, p95 and max |
| Questions written by the builder | You designed for them without noticing | Include unanswerable and nonsense cases |
| One run per question | Hides non-determinism entirely | Three runs minimum |
| Judging quality by reading answers | Fluent prose reads as correct | Mechanical checks: groundedness, source validity, exact figures |
| Warm cache during the run | Measures the cache, not the agent | Clear it, or report hit and miss paths separately |
| Changing two things at once | Cannot attribute the difference | One variable per benchmark run |
| Ignoring failed runs | Excluding errors flatters latency | Report error count alongside every latency figure |
| Using a model to grade its own output | Correlates with fluency, not truth | Model graders for style only; mechanical checks for facts |
The last row is worth expanding, because model-as-judge is popular and quietly unreliable for factual work. A judge model asked "is this answer correct?" without access to the sources is comparing the answer against its own knowledge, which is exactly what you built the agent to avoid depending on. If you use a judge, give it the tool observations and ask a narrow, checkable question — "does every figure in this answer appear in these observations? List any that do not" — rather than a global verdict.
What to do with the results
Run the benchmark before every change and store the report as JSON with a git commit hash. That file is the only thing that can tell you whether last Tuesday's prompt edit helped, and without it you will re-litigate the same argument every month from memory.
Fix in this order, because the categories differ enormously in how much damage they do:
- Groundedness below 1.0. Invented figures are the worst outcome a research agent can produce, because they are indistinguishable from good ones.
- Nonsense and unanswerable categories. An agent that cannot say "I don't know" will confidently mislead somebody.
- The p95 latency tail. Not the mean. The tail is what users complain about and what breaks your platform timeout.
- Tool accuracy. Usually fixed in a docstring, not in code — and it is the cheapest fix on this list.
- Cost. Only once the above are right; a cheap wrong answer has no value.
And be honest about what your numbers can support. Twenty-seven runs tell you whether something is badly broken. They do not tell you that version B is 3 points better than version A. The benchmark's job in a small team is to catch regressions loudly and quickly, and to make one specific claim defensible: that every figure in the agent's output came from somewhere real. That claim is worth more than any accuracy percentage, because it is the one a user can rely on without checking.