AI Agent Fundamentals

Testing, Debugging, and Evaluating Agents


The test suite was green. Fourteen tests, all passing. It took 3 minutes 40 seconds to run and cost about 11 dollars in API calls, so people ran it once a day.

The agent shipped on a Tuesday. By Thursday it had produced a run with 312 model calls, a run that told a customer their refund was processed when it was not, and eleven runs that ended with "I was unable to complete this task" on questions the agent could obviously answer.

None of those were caught, because all fourteen tests did the same thing: send a question to the real model, let the agent run, assert that the final answer contained an expected substring. That tests one path through a stochastic system. It says nothing about what happens when a tool times out, when the model emits malformed output, when an observation comes back empty, or when the agent is 40 steps deep.

Testing an agent is not testing a function. A function has one input and one output; an agent has a trajectory, and nearly everything that goes wrong goes wrong in the middle of it.

What the fourteen green tests were not testingReal model in the loop• Costs 11 dollars a run• Slow, so it runs once a day• Failures are not reproducible• Green means it passed onceScripted model double• Free, and runs in seconds• Same trace every time• Replay a logged failure exactly• Real model reserved for evaluation
Determinism belongs in the tests and stochasticity in the evaluation — mixing them gives you neither.

The four failure modes

1. Repetition loops

Text
Step 7   check_tracking[id=1Z994A] -> "in transit, no update since Tue"Step 8   check_tracking[id=1Z994A] -> "in transit, no update since Tue"...Step 312 check_tracking[id=1Z994A] -> "in transit, no update since Tue"

Cause: an observation that carries no new information. The model's best next guess given an unchanged context is the same guess it made last time.

Detect: hash (tool, args) per run; any count above 1 is a defect.

Fix: feed the repetition back as an observation, not as an exception — "you already ran this and got the same result; try something else or finish". And fix the tool so its observation says something actionable.

Cost of not catching it: traces grow with each step, so cost is roughly quadratic in step count. A 312-step run does not cost 30× a 10-step run; it costs closer to 900×.

2. Hallucinated tool calls

Text
Action: get_customer_lifetime_value[customer_id=8842]Observation: unknown tool 'get_customer_lifetime_value'Action: fetch_ltv[id=8842]Observation: unknown tool 'fetch_ltv'

Cause: the task needs a capability that does not exist, and the model invents a plausible name rather than concluding it cannot proceed.

Detect: count unknown_tool observations per run. Any non-zero rate needs investigating.

Fix: list the available tools in the error, and state explicitly in the prompt that if no tool can supply something, the agent must say so rather than guess. A high rate for one invented name is a product signal — build that tool.

3. Malformed output

Text
Model: I'll check the order status for you now.       Let me look up ORD-4471 in the system.Parser: no Action found

Cause: the model narrates instead of emitting the format. More common with long traces, when the format instructions are far away and buried.

Detect: parse-failure rate per run.

Fix: return the exact required format as the observation; move format rules nearer the end of the prompt; use native tool-calling APIs where available, which move the problem into the provider's validated schema.

4. Wrong termination

PrematureMissed
SymptomAnswers before it has the factsKeeps working after it has them
Looks likeConfident, short, wrongStep budget exhausted
CauseNo goal test; "enough" is a vibeGoal test cannot be satisfied by any state
DetectCompare answer facts against observationsBudget-exhaustion rate
FixMechanical goal test; provenance check on every figureVerify the goal test can return true; add explicit give-up

Premature termination is the dangerous one, because the output looks like success. The refund that was reported as processed and was not came from here: the model saw a 202 Accepted, treated "accepted" as "settled", and stopped.

An agent that stops too early produces a wrong answer that looks right. An agent that stops too late produces a bill. Only one of them is visible without looking.

Debugging

Read the trace, from the failure backwards

When a run goes wrong, find the first step whose thought is wrong, then look at the observation immediately above it. That observation is the cause about four times in five.

Text
Step 3  Action: search_orders[customer=8842, month=July]        Observation: []                              <-- the causeStep 4  Thought: The customer has no July orders. I can answer.        Final Answer: You placed no orders in July.  <-- the symptom

The bug is not in step 4. Given [], step 4's reasoning is correct. The bug is that search_orders returned a bare empty list for a call whose month=July argument it silently ignored, expecting 2026-07. The fix is in the tool:

Text
Observation: {"count": 0, "note": "the 'month' argument must be  YYYY-MM; 'July' was not recognised and no month filter was  applied. This customer has 7 orders overall."}

Now step 4 cannot make that mistake.

Log so that replay is possible

Python
import json, time, hashlibdef log_step(run_id, step, phase, **fields):    print(json.dumps({        "ts": time.time(), "run_id": run_id, "step": step,        "phase": phase, **fields    }))# one line per phase, per steplog_step(run_id, 3, "prompt",  chars=len(prompt),         prompt_sha=hashlib.sha256(prompt.encode()).hexdigest()[:12])log_step(run_id, 3, "llm_out", text=out, in_tok=usage.input_tokens,         out_tok=usage.output_tokens, ms=elapsed_ms)log_step(run_id, 3, "parsed",  ok=True, tool="search_orders",         args={"customer": 8842, "month": "July"})log_step(run_id, 3, "exec",    ok=True, ms=41, result_chars=2)log_step(run_id, 3, "obs",     text="[]")

Structured JSON lines, one per phase, with run_id and step on every line. That shape lets you answer operational questions with a one-liner:

Bash
# which tools fail most?jq -r 'select(.phase=="exec" and .ok==false) | .tool' logs.jsonl \  | sort | uniq -c | sort -rn# runs with repeated identical callsjq -r 'select(.phase=="parsed") | "\(.run_id) \(.tool)\(.args)"' \  logs.jsonl | sort | uniq -c | awk '$1 > 1'# p95 step countjq -r 'select(.phase=="obs") | .run_id' logs.jsonl | sort | uniq -c \  | awk '{print $1}' | sort -n | awk '{a[NR]=$1} END {print a[int(NR*0.95)]}'

Log the prompt's hash rather than the prompt itself. Full prompts are enormous and mostly identical; a hash tells you whether the stable prefix changed, which is what you actually want to know.

Replay

Because the loop is a pure function of (prompt, model output, tool result), a recorded run can be replayed exactly:

Python
class ReplayClient(LLMClient):    """Returns the exact model outputs a recorded run produced."""    def __init__(self, log_path, run_id):        self.outs = [json.loads(l)["text"]                     for l in open(log_path)                     if json.loads(l).get("run_id") == run_id                     and json.loads(l)["phase"] == "llm_out"]        self.i = 0    def complete(self, prompt, stop=None):        out = self.outs[self.i]; self.i += 1        return out

Replay lets you change the parser, the trace compaction, the budgets or the tool wrappers and see immediately whether the historical failure still happens — deterministically, instantly, for free. Any production failure worth fixing should become a replay fixture.

Unit tests

Why the real model cannot be in the loop

Real modelScripted stand-in
DeterminismNone, even at temperature 0Total
Speed1–5s per stepMicroseconds
Cost of a 200-test suiteTens of dollars per runZero
Can you test a timeout?NoYes
Can you test malformed output?Only by luckYes, exactly
Can you test step 40?Slowly and expensivelyInstantly
FlakesConstantlyNever

The decisive row is "can you test malformed output". You cannot reliably make a real model produce a specific broken response, and broken responses are precisely what your loop must survive. Scripting the model turns "hope this never happens" into a test case.

Scripting a conversation

Python
import pytestdef make_agent(turns, tools=None):    return Agent(ScriptedClient(turns), tools or registry, max_steps=10)def test_happy_path_two_tools():    agent = make_agent([        "Thought: Need the order.\nAction: get_order[order_id=ORD-4471]",        "Thought: Delivered, so 30 days.\n"        "Action: get_refund_policy[status=delivered]",        "Thought: I have both facts.\nFinal Answer: 30 days, from "        "2026-07-12.",    ])    r = agent.run("Refund window for ORD-4471?")    assert r["status"] == "ok"    assert r["steps"] == 3    assert "30 days" in r["answer"]def test_recovers_from_bad_parameter_name():    agent = make_agent([        "Thought: Look it up.\nAction: get_order[id=ORD-4471]",        "Thought: Wrong parameter name.\n"        "Action: get_order[order_id=ORD-4471]",        "Thought: Got it.\nFinal Answer: delivered.",    ])    r = agent.run("Status of ORD-4471?")    assert r["status"] == "ok"    obs = r["trace"].entries[0]["observation"]    assert "order_id" in obs           # error names the correctiondef test_unparseable_output_does_not_crash():    agent = make_agent([        "I'll just have a look at that order for you.",   # no Action        "Thought: Formatting properly.\n"        "Action: get_order[order_id=ORD-4471]",        "Thought: Done.\nFinal Answer: delivered.",    ])    r = agent.run("Status?")    assert r["status"] == "ok"    assert "Action:" in r["trace"].entries[0]["observation"]def test_repeat_detection_breaks_a_loop():    same = "Thought: Checking.\nAction: get_order[order_id=ORD-4471]"    agent = make_agent([same] * 10)    r = agent.run("Status?")    assert r["status"] == "max_steps"    later = r["trace"].entries[3]["observation"]    assert "already ran" in laterdef test_tool_exception_becomes_an_observation():    reg = Registry()    @reg.register    def flaky(x: str) -> str:        """Always fails."""        raise ConnectionError("upstream refused")    agent = make_agent([        "Thought: Trying.\nAction: flaky[x=1]",        "Thought: It failed; I will report that.\n"        "Final Answer: The service is unavailable.",    ], tools=reg)    r = agent.run("Do the thing")    assert r["status"] == "ok"    assert "ConnectionError" in r["trace"].entries[0]["observation"]def test_unknown_tool_lists_the_real_ones():    agent = make_agent([        "Thought: Trying.\nAction: get_ltv[customer_id=8842]",        "Thought: No such tool.\nFinal Answer: I cannot look that up.",    ])    r = agent.run("What is customer 8842 worth?")    assert "get_order" in r["trace"].entries[0]["observation"]def test_step_budget_is_enforced():    agent = make_agent(        ["Thought: n.\nAction: get_order[order_id=ORD-%d]" % i         for i in range(20)])    r = agent.run("go")    assert r["status"] == "max_steps"    assert r["steps"] == 10@pytest.mark.parametrize("text,expect", [    ("Action: t[a=1]",              ("action", "t", {"a": 1})),    ("Action: t[a=1, b=hello]",     ("action", "t", {"a": 1, "b": "hello"})),    ("Action: t[a='x, y']",         ("action", "t", {"a": "x, y"})),    ("Action:  t [a=1] ",           ("action", "t", {"a": 1})),    ("Action: t[]",                 ("action", "t", {})),    ("Final Answer: 42",            ("final", "42")),    ("just some prose",             ("error",)),])def test_parser(text, expect):    assert parse(text)[:len(expect)] == expect

Eight tests, no network, runs in well under a second, and between them they cover every one of the four failure modes. Compare that with fourteen end-to-end tests that cost 11 dollars and check one substring.

The layers, and what each is for

LayerTestsModelToolsCountRuntime
UnitParser, registry, trace, budgetsNoneFake50–200Under a second
LoopRecovery, loops, terminationScriptedFake20–60Seconds
IntegrationReal tools against a test databaseScriptedReal10–30Tens of seconds
EvaluationEnd-to-end quality on a task setRealReal100–500 casesMinutes, costs money

The first three gate every commit. The fourth runs nightly or before a release — it is a measurement, not a gate, because it is stochastic and a single failure does not mean a regression.

Evaluation

The metrics that matter

MetricDefinitionTargetReveals
Task success rateRuns meeting the goal test ÷ all runsDomain-specificOverall quality
Tool-choice accuracyCorrect tool ÷ all callsAbove 95%Tool description quality
Step efficiencyActual steps ÷ minimum neededUnder 1.5Wandering
Termination correctnessStopped at the right moment ÷ all runsAbove 95%Goal-test quality
Cost per taskMean and p95 in centsBudget-specificWhether it is viable at volume
Recovery rateRuns succeeding after an error ÷ runs with an errorAbove 80%Error-message quality

A worked evaluation

200 labelled cases, one run each.

Text
Successes          148 / 200  = 74.0%Median steps         4p95 steps           17Mean cost         6.2 centsp95 cost         31.0 centsFailures (52), by first defect in the trace:   wrong tool chosen          21   (40.4% of failures)   repetition loop            14   (26.9%)   premature termination       9   (17.3%)   parse failure               8   (15.4%)

Read it as a causal chain rather than four separate problems. Wrong tool choices produce useless observations. Useless observations give the model nothing new, which produces repetition loops. Repetition burns the step budget, which pushes p95 steps to 17 and p95 cost to 31 cents — five times the mean.

So the intervention is at the head of the chain: rewrite the descriptions of the tools involved in those 21 cases. Suppose that recovers 15 of the 21 wrong-tool failures and, because fewer runs now get useless observations, half the loops as well:

Text
148 + 15 + 7  =  170 / 200  =  85.0%       (up 11 points)

Now price it. At 40,000 runs a month:

Text
before:  40,000 x 6.2 cents  =  2,480 USD / monthafter :  40,000 x 4.8 cents  =  1,920 USD / month                     saving  =    560 USD / month

The mean falls to about 4.8 cents because the expensive runs are the looping ones, and halving the loops removes most of the tail. An afternoon spent on tool descriptions bought 11 points of success rate and 560 dollars a month. No prompt tuning, no model change.

Agent metrics are rarely independent. When four numbers look bad at once, find the earliest link in the chain and fix that one — the other three usually follow.

Non-determinism, and what to do about it

Even at temperature 0 the same input can yield different outputs — and many current reasoning models do not accept a temperature setting at all — so a single run of a case is a sample, not a measurement. Two consequences.

Report pass@k honestly. If a case succeeds on 2 of 3 attempts, pass@1 is 67% and pass@3 (succeeds at least once) is 100%. Which one you quote depends on your product: pass@1 for a system that answers once, pass@3 only where a user can genuinely retry.

And do the significance arithmetic before believing an improvement. On 200 cases, a success rate near 75% has a standard error of about

0.75×0.25200≈0.031\sqrt{\frac{0.75 \times 0.25}{200}} \approx 0.031

— roughly 3 percentage points, so a 95% interval is about ±6 points. A change from 74% to 77% is noise. A change from 74% to 85% is real. Chasing three-point movements on a 200-case set is how teams spend a fortnight shipping nothing.

Judging correctness

MethodCostReliabilityUse for
Exact matchFreePerfect where it appliesNumbers, IDs, classifications
Required substringsFreeHighAnswers that must contain specific facts
Structural checkFreeHighSource counts, field presence, formats
Provenance grepFreeHighEvery figure appears in some observation
Golden trajectoryFreeBrittle — many valid paths existRegression only, never absolute quality
Model as judgeOne call/caseModerate; biased towards fluencyOpen-ended answers, with a rubric
Human reviewMinutes/caseHighestCalibrating the automatic judges

Two cautions. A model judge shown the agent's reasoning tends to endorse it, so give the judge the question, the answer and the raw observations — never the chain of thought. And judges score fluent wrong answers generously, which is exactly the failure you most need to catch, so calibrate against 30 human-labelled cases before trusting the number.

Where people get this wrong

Testing only the happy path. The happy path is the one case that mostly works already. Every test worth writing is about what happens when something breaks.

Real API calls in unit tests. Slow, expensive, flaky, and unable to test the cases you care about most.

Asserting on exact final wording. The model rephrases; the test breaks; someone deletes the test. Assert on facts, structure and trajectory properties instead.

Golden trajectories as a quality gate. Many different action sequences are equally correct. Use them to detect change, not to define correctness.

Reporting only the mean cost. The mean hides the loops. The p95 is where your money goes and where your incidents live.

Chasing noise. Without an error bar you cannot tell an improvement from a coin flip.

Discarding traces from successful runs. A run that succeeded in 14 steps when 4 would do is a defect you cannot see if you only keep failures.

What this means when you build one

Make the model swappable on the first day. Everything in this lesson depends on being able to hand the agent a scripted sequence of turns instead of a network call, and retrofitting that into a loop that calls the API inline is an unpleasant refactor. One interface, two implementations, done in ten minutes at the start.

Then write your first tests about failure. Malformed output, unknown tool, tool exception, identical repeated call, budget exhaustion. Those five tests catch more real defects than fifty happy-path assertions, and they run in under a second so nobody is tempted to skip them.

Log every run as structured JSON lines with a run ID and a step number, and keep enough to replay. The first time a production failure turns into a replayed fixture that you fix in ten minutes without spending a rupee on API calls, the logging pays for itself permanently.

Build an evaluation set of 100 to 200 real cases with mechanical goal tests, and when you look at the results, resist fixing the metric that looks worst. Trace the chain backwards to the earliest defect. In the numbers above, the visible symptom was a 17-step p95 and a 31-cent tail — and the actual bug was two tool descriptions a person wrote in four seconds.