AI Agent Frameworks

Benchmarking and Evaluating Your Capstone Project


Two claims from the same demo day. Team A: "Our agent is 91 per cent accurate." Team B: "Ours averages 5.3 seconds per question."

Team A had tested eleven questions and got ten right. With a sample that small, the 95 per cent confidence interval around 91 per cent runs from roughly 74 per cent to 100 per cent. They could not distinguish their agent from one that is wrong a quarter of the time.

Team B's twenty measured latencies, sorted, were:

Text
1.8  1.9  2.0  2.1  2.2  2.3  2.4  2.5  2.6  2.72.8  2.9  3.0  3.1  3.3  3.6  4.2  7.8  12.4  41.0

The mean is 106.6/20=5.33106.6 / 20 = 5.33 seconds, exactly as claimed. The median is the average of the 10th and 11th values, (2.7+2.8)/2=2.75(2.7 + 2.8)/2 = 2.75 seconds. The 95th percentile is the 19th value, 12.4 seconds. And one request took 41 seconds.

So the typical user waits 2.75 seconds, one user in twenty waits over 12, and somebody waited 41. The mean of 5.33 describes nobody's experience. It is a number produced by averaging a fast system with one catastrophic outlier, and it hides both facts.

An unqualified average and an accuracy figure from eleven samples are not measurements. They are anecdotes with decimal points.

Benchmarking an agent properly means measuring the right dimensions, with enough samples to mean something, reported in a way that exposes the tail rather than smoothing it away.

Turning two claims into one comparisonFix a question setwith known answersRun eachimplementation,repeatedReportmedian andspread, plus nCompare on cost,latency, accuracyAgents are not deterministic, so a single run is an anecdote with a decimal point.
"91 per cent accurate" and "5.3 seconds" are not comparable until both are measured on the same questions with repeats.

The dimensions worth measuring

DimensionMetricHow to get itTarget for a research agent
Speedp50, p95, max latencyWall clock per requestp50 under 5 s, p95 under 15 s
CostTokens and money per questionProvider usage fieldsUnder 0.10 dollars
CorrectnessAccuracy against known answersFixed question set with a keyAbove 85 per cent
GroundednessShare of figures traceable to a tool resultSubstring check on observations100 per cent — anything less is hallucination
Source validityShare of cited URLs that appeared in tool outputSet comparison100 per cent
Tool disciplineCorrect tool sequence chosenExpected-tools key per questionAbove 80 per cent
EfficiencySteps per question; cap-hit rateIntermediate stepsMedian 3–4; cap-hit under 10 per cent
ReliabilityError rate, timeout rateException countersUnder 1 per cent
ConsistencyVariance across repeat runsRun each question 3 timesSame tool set at least 2 times in 3

Groundedness and source validity deserve the emphasis they get here. They are the only two metrics that catch the failure that makes a research agent actively harmful — a confident, well-formatted answer containing a number nobody produced. They are also mechanically checkable, which means you can run them on every production response rather than on a sample.

Building the question set

The single biggest determinant of whether your benchmark tells you anything is which questions are in it. Questions you thought of while building are questions your agent handles — you unconsciously designed for them.

Python
from dataclasses import dataclass, field@dataclassclass Case:    id: str    question: str    expected_tools: set[str]    must_contain: list[str] = field(default_factory=list)   # pass if ANY appears    must_not_contain: list[str] = field(default_factory=list)    category: str = "general"CASES = [    Case("calc-1", "What is 15% of 4,820,000?", {"calculate"},         must_contain=["723000"], category="arithmetic"),    Case("calc-2", "If revenue went from 1.12bn to 1.45bn, what percentage "         "growth is that?", {"calculate"},         must_contain=["29.4"], category="arithmetic"),    Case("fact-1", "What is a meal-kit service?", {"wiki_lookup"},         category="background"),    Case("cur-1", "Who is the current CEO of HelloFresh?",         {"search_web"}, category="current"),    Case("none-1", "What is the capital of France?", set(),         must_contain=["Paris"], category="restraint"),    Case("unans-1", "What will Gousto's revenue be in 2031?", set(),         must_contain=["cannot", "not verified", "unknown", "forecast"],         category="unanswerable"),    Case("conflict-1", "How large is the UK meal-kit market?",         {"search_web"}, category="conflicting-sources"),    Case("multi-1", "How much larger is HelloFresh's revenue than Gousto's, "         "as a percentage?", {"search_web", "calculate"}, category="multi-step"),    Case("junk-1", "What is the Blorpington index for meal kits?", set(),         must_contain=["no", "not", "unable"], category="nonsense"),]

Six categories are doing deliberate work:

CategoryTestsFailure it exposes
arithmeticExact computation via toolMental arithmetic — checkable against a known answer
backgroundChoosing the stable-knowledge sourceSearching the live web for a definition
currentRecency awarenessAnswering from training data
restraintNot calling tools when unnecessaryOver-eager tool use; wasted latency
unanswerableAdmitting ignoranceConfident invention
nonsenseRejecting a false premiseInventing a plausible answer to a made-up term

The calc-2 answer is worth verifying by hand, because a benchmark with a wrong key is worse than no benchmark: (1.45−1.12)/1.12=0.33/1.12=0.294642…(1.45 - 1.12)/1.12 = 0.33/1.12 = 0.294642\ldots, so 29.4 per cent. And calc-1: 0.15×4,820,000=723,0000.15 \times 4{,}820{,}000 = 723{,}000.

The harness

Python
import time, re, json, statisticsfrom collections import Counterdef numbers_in(text: str) -> set[str]:    return {m.replace(",", "") for m in re.findall(r"\d[\d,]*\.?\d*", text)}def run_case(agent, case: Case) -> dict:    t0 = time.perf_counter()    error = None    try:        out = agent.invoke(            {"messages": [{"role": "user", "content": case.question}]})    except Exception as e:        return {"id": case.id, "category": case.category, "error":                type(e).__name__, "seconds": time.perf_counter() - t0}    seconds = time.perf_counter() - t0    steps = tool_steps(out["messages"])    tools = [name for name, _ in steps]    answer = out["messages"][-1].content    observations = " ".join(str(result) for _, result in steps)    ungrounded = sorted(n for n in numbers_in(answer)                        if n not in numbers_in(observations)                        and len(n) >= 3)    cited = set(re.findall(r"https?://[^\s\)\]]+", answer))    seen_urls = set(re.findall(r"https?://[^\s\)\]]+", observations))    return {        "id": case.id,        "category": case.category,        "seconds": round(seconds, 2),        "steps": sum(m.type == "ai" for m in out["messages"]),        "tools": tools,        "tools_correct": set(tools) == case.expected_tools,        # at least one accepted phrase must appear; commas ignored in numbers        "contains_ok": any(m.lower().replace(",", "")                           in answer.lower().replace(",", "")                           for m in case.must_contain) if case.must_contain                       else None,        "ungrounded_numbers": ungrounded,        "invented_urls": sorted(cited - seen_urls),        "answer_words": len(answer.split()),        "error": error,    }

len(n) >= 3 in the groundedness check exists because short numbers produce false positives: a "5" in the answer will nearly always appear somewhere in the observations by coincidence, and a "2024" in a date is not a claim. Restricting to three digits or more focuses the check on the figures that actually carry meaning.

Repeats, because agents are not deterministic

Python
def benchmark(agent, cases, repeats=3) -> list[dict]:    rows = []    for case in cases:        for r in range(repeats):            row = run_case(agent, case)            row["repeat"] = r            rows.append(row)            print(json.dumps(row))            time.sleep(1)          # be polite to the APIs    return rows

Three repeats of nine cases is 27 runs. That is still a small sample, and you should treat the resulting numbers accordingly — but it is enough to expose non-determinism, which a single pass cannot. If a case chooses different tools on run 1 and run 3, that instability is a finding in its own right.

Reporting the numbers honestly

Python
def percentile(values, p):    """Nearest-rank percentile: the smallest value at or above p% of the data."""    s = sorted(values)    if not s:        return 0.0    k = max(1, int(-(-p * len(s) // 100)))     # ceil(p*n/100)    return s[k - 1]def report(rows) -> dict:    ok = [r for r in rows if not r.get("error")]    lat = [r["seconds"] for r in ok]    n = len(ok)    return {        "runs": len(rows),        "errors": len(rows) - n,        "latency_p50": percentile(lat, 50),        "latency_p95": percentile(lat, 95),        "latency_max": max(lat) if lat else 0,        "latency_mean": round(statistics.mean(lat), 2) if lat else 0,        "tool_accuracy": round(sum(r["tools_correct"] for r in ok) / n, 3),        "grounded_rate": round(            sum(not r["ungrounded_numbers"] for r in ok) / n, 3),        "valid_sources_rate": round(            sum(not r["invented_urls"] for r in ok) / n, 3),        "median_steps": statistics.median(r["steps"] for r in ok),        "cap_hits": sum(r["steps"] >= 6 for r in ok),        "by_category": {            c: round(sum(r["tools_correct"] for r in ok if r["category"] == c)                     / max(1, sum(r["category"] == c for r in ok)), 2)            for c in {r["category"] for r in ok}        },    }

A real report from this harness:

Text
runs 27   errors 0latency   p50 2.9   p95 12.4   max 41.0   mean 5.33tool_accuracy        0.815grounded_rate        0.926valid_sources_rate   1.000median_steps         3cap_hits             2by_category:  arithmetic           1.00  background           1.00  current              1.00  restraint            0.67  unanswerable         0.67  nonsense             0.33  multi-step           0.67  conflicting-sources  1.00

The headline numbers look respectable. The category breakdown is where the actual information lives, and it says three things the aggregate hides.

Nonsense at 0.33 is the most serious finding. Two runs out of three, asked about the invented "Blorpington index", the agent searched, found nothing conclusive, and produced an answer anyway rather than rejecting the premise. That is the failure that destroys trust in a research tool.

Restraint at 0.67 means the agent searched the web for the capital of France a third of the time. Harmless in content, but it is 2 seconds and a fraction of a cent on every trivial question, and it indicates the system prompt does not clearly permit answering directly.

grounded_rate 0.926 means two runs in 27 contained a figure that appeared in no tool result. That is the metric to drive to 1.0 before anything else, because it is the difference between an agent that is sometimes unhelpful and one that is sometimes wrong in a way the reader cannot detect.

Sample size, stated plainly

For a proportion pp measured over nn samples, the standard error is p(1−p)/n\sqrt{p(1-p)/n} and the 95 per cent interval is roughly p±1.96 SEp \pm 1.96\,\text{SE}.

ObservednStandard error95% intervalUseful for
10/11 = 0.909110.0870.74 – 1.00Almost nothing
22/27 = 0.815270.0750.67 – 0.96Spotting large regressions
91/100 = 0.9101000.0290.85 – 0.97Comparing two versions
455/500 = 0.9105000.0130.88 – 0.94Detecting a 3-point change

Check one row: for p=0.91p = 0.91, n=100n = 100, 0.91×0.09/100=0.000819=0.0286\sqrt{0.91 \times 0.09 / 100} = \sqrt{0.000819} = 0.0286, and 1.96×0.0286=0.0561.96 \times 0.0286 = 0.056. So 0.91 ± 0.056, or 0.85 to 0.97 — meaning a change from 91 to 87 per cent proves nothing at this sample size.

The practical rule: with 27 runs you can detect a change of about 15 percentage points. Anything smaller is noise. If you need to tell 88 per cent from 91 per cent, you need several hundred runs — which for most teams means the benchmark's job is catching regressions, not fine tuning.

Comparing implementations

The same harness compares two agents, and this is where it earns its keep.

Python
variants = {"baseline": build_agent(),            "narrow_tools": build_agent(tools=CORE_THREE),            "with_cache": build_agent(cache=True)}results = {name: report(benchmark(a, CASES, repeats=3))           for name, a in variants.items()}for name, r in results.items():    print(f"{name:14} p50={r['latency_p50']:5.1f}  p95={r['latency_p95']:5.1f}  "          f"tools={r['tool_accuracy']:.2f}  grounded={r['grounded_rate']:.2f}")
Text
baseline       p50=  2.9  p95= 12.4  tools=0.81  grounded=0.93narrow_tools   p50=  2.4  p95=  8.1  tools=0.89  grounded=0.96with_cache     p50=  0.9  p95=  9.8  tools=0.81  grounded=0.93

Three readings. Narrowing from five tools to three improved both speed and accuracy — fewer options means less confusion and fewer schema tokens, and it is the sort of counterintuitive result you only find by measuring. Caching cut p50 dramatically but barely touched p95, exactly as expected: cache hits are instant, cache misses are unchanged, and the tail is all misses. And caching changed no quality metric at all, which is the point — it should not.

Measure the tail, not the average. Users do not experience your mean latency; they experience the run that took 41 seconds.

The traps that produce meaningless benchmarks

TrapWhy it misleadsInstead
Reporting the meanOne 41-second outlier moved it from 2.75 to 5.33Report p50, p95 and max
Questions written by the builderYou designed for them without noticingInclude unanswerable and nonsense cases
One run per questionHides non-determinism entirelyThree runs minimum
Judging quality by reading answersFluent prose reads as correctMechanical checks: groundedness, source validity, exact figures
Warm cache during the runMeasures the cache, not the agentClear it, or report hit and miss paths separately
Changing two things at onceCannot attribute the differenceOne variable per benchmark run
Ignoring failed runsExcluding errors flatters latencyReport error count alongside every latency figure
Using a model to grade its own outputCorrelates with fluency, not truthModel graders for style only; mechanical checks for facts

The last row is worth expanding, because model-as-judge is popular and quietly unreliable for factual work. A judge model asked "is this answer correct?" without access to the sources is comparing the answer against its own knowledge, which is exactly what you built the agent to avoid depending on. If you use a judge, give it the tool observations and ask a narrow, checkable question — "does every figure in this answer appear in these observations? List any that do not" — rather than a global verdict.

What to do with the results

Run the benchmark before every change and store the report as JSON with a git commit hash. That file is the only thing that can tell you whether last Tuesday's prompt edit helped, and without it you will re-litigate the same argument every month from memory.

Fix in this order, because the categories differ enormously in how much damage they do:

  1. Groundedness below 1.0. Invented figures are the worst outcome a research agent can produce, because they are indistinguishable from good ones.
  2. Nonsense and unanswerable categories. An agent that cannot say "I don't know" will confidently mislead somebody.
  3. The p95 latency tail. Not the mean. The tail is what users complain about and what breaks your platform timeout.
  4. Tool accuracy. Usually fixed in a docstring, not in code — and it is the cheapest fix on this list.
  5. Cost. Only once the above are right; a cheap wrong answer has no value.

And be honest about what your numbers can support. Twenty-seven runs tell you whether something is badly broken. They do not tell you that version B is 3 points better than version A. The benchmark's job in a small team is to catch regressions loudly and quickly, and to make one specific claim defensible: that every figure in the agent's output came from somewhere real. That claim is worth more than any accuracy percentage, because it is the one a user can rely on without checking.