Course Content
Model Context Protocol (MCP)
3 sections · 8 lessons
Testing and Reliability for MCP Agents
A team has 340 tests for their MCP agent. All green, on every commit, for six weeks. In production the agent fails roughly one run in six: sometimes it hangs for two minutes and returns nothing, sometimes it answers confidently with no sources, and once it told a customer their refund window was 60 days when the policy had said 30 days since July.
Read the suite and the reason is immediate. Every test patches the client manager:
manager.call = AsyncMock(return_value="[POL-114] Refunds within 30 days.")That line asserts that if a tool returns a well-formed string, the agent handles it. Not one test ever sent a real JSON-RPC request, sent an argument through a real schema validator, saw an HTTP 429, waited out a timeout, or watched what happens when the tool returns 40,000 tokens. The suite tests the glue between mocks. The production failures all live in the parts that were mocked away.
Testing an agent is not harder than testing ordinary software, but the boundaries are in different places, and putting them in the usual places produces exactly this: high coverage and low confidence.
Why the pyramid bends
The critical observation is that an agent system has one stochastic component and many deterministic ones, and they need completely different treatment.
| Layer | Deterministic? | How to test it | Assert on |
|---|---|---|---|
| Server handlers | Yes | Unit tests, direct calls | Exact outputs |
| Tool schemas | Yes | Property tests over the catalogue | Structural rules |
| Transport and discovery | Yes | Integration, real subprocess | Server answers, tools list |
| Retry / breaker / timeout | Yes, with a fake clock | Fault injection | Number and timing of attempts |
| Model tool selection | No | Repeated runs against fixtures | Rates and invariants, never exact text |
| Final answer wording | No | Property checks, spot review | Citations present, no unsupported claims |
Assert exact behaviour everywhere the system is deterministic, and assert only properties and rates where it is not. Mixing those up gives you either a flaky suite or a useless one.
Unit-testing the server
The MCP Python SDK can run a client and server in the same process: pass the server object itself to Client and it connects over an in-memory transport. This is the sweet spot: no subprocess, no network, but real protocol.
1import pytest2from mcp import Client3from server import mcp # the MCPServer instance under test45@pytest.mark.anyio6async def test_tool_is_advertised_with_a_usable_schema():7 async with Client(mcp) as client:8 tools = {t.name: t for t in (await client.list_tools()).tools}9 assert "search_orders" in tools10 schema = tools["search_orders"].input_schema11 assert schema["required"] == ["customer_email"]12 assert schema["properties"]["limit"]["maximum"] == 5013 assert "shipped" in schema["properties"]["status"]["anyOf"][0]["enum"]1415@pytest.mark.anyio16async def test_bad_date_returns_a_correctable_message_not_an_exception():17 async with Client(mcp) as client:18 result = await client.call_tool(19 "search_orders", {"customer_email": "a@b.com", "since": "last week"})20 text = result.content[0].text21 assert "ISO-8601" in text # tells the model the format22 assert "2026-" in text # and shows an exampleThat second test encodes a rule most suites miss: an invalid argument must produce a message the model can act on. Asserting only that the call did not crash lets someone "fix" a bug by returning the string "error", which passes the test and destroys the agent's ability to recover.
Bounds are the tests people skip
1@pytest.mark.anyio2async def test_result_is_bounded_even_when_the_data_is_not(seeded_db_with_4300_orders):3 async with Client(mcp) as client:4 result = await client.call_tool(5 "search_orders", {"customer_email": "bulk@example.com", "limit": 50})6 text = result.content[0].text7 assert len(text) <= 40_0008 assert text.count("\n") <= 55 # 50 rows plus header and notice9 assert "more matched" in text # the model is told it was truncatedA truncation that does not announce itself is worse than no truncation, because the model treats a partial result as complete and reasons from it.
Testing the catalogue as data
Some of the highest-value tests iterate over every tool and check structural rules. They cost twenty lines and catch a whole class of mistakes at the moment they are introduced.
1import re2from jsonschema import Draft202012Validator34NAME_RE = re.compile(r"^[a-z][a-z0-9_]{2,63}$")56@pytest.mark.anyio7async def test_every_tool_definition_is_well_formed():8 async with Client(mcp) as client:9 tools = (await client.list_tools()).tools1011 assert len({t.name for t in tools}) == len(tools), "duplicate tool names"1213 for t in tools:14 assert NAME_RE.match(t.name), f"bad tool name: {t.name}"15 Draft202012Validator.check_schema(t.input_schema) # the schema itself is valid1617 desc = (t.description or "").strip()18 assert len(desc) >= 40, f"{t.name}: description too thin to choose from"19 assert len(desc) <= 1024, f"{t.name}: description will bloat every prompt"2021 props = t.input_schema.get("properties", {})22 for field, spec in props.items():23 assert spec.get("description"), f"{t.name}.{field} has no description"24 if spec.get("type") == "integer":25 assert "maximum" in spec, f"{t.name}.{field} is an unbounded integer"2627 total = sum(len(t.name) + len(t.description or "") +28 len(str(t.input_schema)) for t in tools) // 429 assert total < 6_000, f"tool catalogue is ~{total} tokens per model call"That last assertion is a budget guard in test form. A catalogue that quietly grows from 4,200 to 16,000 tokens raises the cost of every model call in the system, and nothing else in a normal CI pipeline notices.
Integration: use the real transport
In-memory connections skip subprocess launch, environment handling and stream framing — which is where the "works on my machine" failures live. At least one test must spawn the server for real.
1from mcp import Client, StdioServerParameters23@pytest.mark.anyio4async def test_server_starts_from_a_clean_environment_and_lists_tools():5 params = StdioServerParameters(6 command="uv", args=["run", "python", "server.py"],7 env={"DATABASE_URL": TEST_DSN, "PATH": os.environ["PATH"]}, # nothing else8 )9 async with Client(params) as client:10 listing = await asyncio.wait_for(client.list_tools(), timeout=10)11 assert client.protocol_version == "2026-07-28" # not a legacy fallback12 assert client.server_capabilities.tools is not None13 assert {t.name for t in listing.tools} == EXPECTED_TOOL_NAMESPassing a minimal env is the point of the test. The protocol-version assertion is a cheap guard too: if a dependency pin drags the server back to an SDK that only speaks the old handshake, the client silently falls back, and this line tells you. Servers routinely work locally because the developer's shell exports a variable that the deployment does not, and this test fails loudly on exactly that.
The four failure patterns
Real MCP agents fail in four recognisable ways. Each needs its own mechanism and its own test.
Timeouts
Every outbound call needs a deadline, and the deadline must be tested with a server that genuinely hangs — not with a mock that raises TimeoutError, which proves only that your except clause has correct syntax.
1@pytest.mark.anyio2async def test_hanging_server_is_abandoned_with_a_useful_message():3 manager = await start_manager(MOCK_FAIL="timeout") # the mock sleeps 120 s4 started = time.monotonic()5 text = await manager.call("docs__semantic_search", {"query": "refunds"},6 timeout=2.0)7 elapsed = time.monotonic() - started8 assert 2.0 <= elapsed < 3.0, "timeout was not enforced"9 assert "did not respond" in text10 assert "narrower query" in text # actionable for the modelTimeout budgets must nest. If the agent's overall budget is 180 seconds and a single tool may take 45, then four sequential timeouts exhaust the run before the model writes anything. Set per-call timeouts so that the worst realistic sequence still leaves time to answer: with a 12-iteration cap, 45 seconds each is 540 seconds of potential blocking against a 180-second budget — the numbers do not fit, and 15 seconds per call is the honest ceiling.
Retry with backoff
Test the schedule, not just the outcome. Capture the sleeps.
1@pytest.mark.anyio2async def test_backoff_grows_and_is_jittered(monkeypatch):3 waits: list[float] = []4 async def fake_sleep(s): waits.append(s)5 monkeypatch.setattr(asyncio, "sleep", fake_sleep)67 attempts = 08 async def flaky():9 nonlocal attempts10 attempts += 111 if attempts < 4:12 raise Retryable("429")13 return "ok"1415 assert await call_with_retry(flaky) == "ok"16 assert attempts == 417 assert len(waits) == 318 # Each wait is bounded by the exponential ceiling for its attempt...19 assert waits[0] <= 1.0 and waits[1] <= 2.0 and waits[2] <= 4.020 # ...and jitter means they are not all identical.21 assert len(set(waits)) > 1The final assertion is the one that catches the outage from the opening: without jitter every client waits the identical interval, and the retries arrive as a synchronised burst that re-triggers the throttle.
Circuit breaker
Retries handle a blip. They are the wrong tool for an upstream that is down, because every call still pays the full timeout before failing. Work the numbers: an upstream out for ten minutes, an agent making 20 calls a minute, a 45-second timeout on each — that is 200 calls each blocking a worker for 45 seconds, or 2.5 hours of worker time spent waiting for a service you already know is broken.
A circuit breaker converts that into 5 slow failures followed by instant ones.
1import time23class CircuitBreaker:4 """CLOSED -> (5 failures in 60 s) -> OPEN -> (after cooldown) -> HALF_OPEN"""56 def __init__(self, threshold=5, window=60.0, cooldown=30.0, max_cooldown=300.0):7 self.threshold, self.window = threshold, window8 self.cooldown, self.max_cooldown = cooldown, max_cooldown9 self.state, self.failures, self.opened_at = "CLOSED", [], 0.010 self.current_cooldown = cooldown1112 def allow(self) -> bool:13 now = time.monotonic()14 if self.state == "OPEN":15 if now - self.opened_at >= self.current_cooldown:16 self.state = "HALF_OPEN" # let exactly one trial through17 return True18 return False19 return True2021 def record(self, ok: bool) -> None:22 now = time.monotonic()23 if ok:24 self.state, self.failures = "CLOSED", []25 self.current_cooldown = self.cooldown # reset the escalation26 return27 if self.state == "HALF_OPEN": # trial failed: back off harder28 self.state, self.opened_at = "OPEN", now29 self.current_cooldown = min(self.current_cooldown * 2, self.max_cooldown)30 return31 self.failures = [t for t in self.failures if now - t < self.window] + [now]32 if len(self.failures) >= self.threshold:33 self.state, self.opened_at = "OPEN", now1@pytest.mark.anyio2async def test_breaker_opens_then_probes_then_closes(monkeypatch):3 clock = [1000.0]4 monkeypatch.setattr(time, "monotonic", lambda: clock[0])5 cb = CircuitBreaker(threshold=5, window=60.0, cooldown=30.0)67 for _ in range(5):8 assert cb.allow()9 cb.record(ok=False)10 assert cb.state == "OPEN"11 assert cb.allow() is False # fails instantly, no timeout paid1213 clock[0] += 31.014 assert cb.allow() is True # one probe permitted15 assert cb.state == "HALF_OPEN"16 cb.record(ok=False)17 assert cb.state == "OPEN"18 assert cb.current_cooldown == 60.0 # escalated 30 -> 601920 clock[0] += 61.021 assert cb.allow() is True22 cb.record(ok=True)23 assert cb.state == "CLOSED"Fallback
When the breaker is open, something still has to answer. A fallback chain degrades in a defined order and — critically — tells the caller which tier answered.
1async def search_with_fallback(query: str) -> tuple[str, str]:2 """Returns (text, tier). The tier travels with the result so the model3 can qualify its answer instead of presenting stale data as current."""4 if breakers["docs"].allow():5 try:6 text = await manager.call("docs__semantic_search", {"query": query})7 breakers["docs"].record(ok=True)8 return text, "primary"9 except Exception:10 breakers["docs"].record(ok=False)1112 cached = cache.get(query) # tier 2: possibly stale13 if cached:14 return f"{cached}\n\n(from cache, {cache.age(query)//60} minutes old)", "cache"1516 if breakers["web"].allow(): # tier 3: different source17 try:18 return await manager.call("web__web_search", {"query": query}), "degraded"19 except Exception:20 breakers["web"].record(ok=False)2122 return ("No source is currently reachable for this query. "23 "State that the information could not be verified."), "unavailable"Test that each tier is reachable and that the tier label is truthful. A fallback that silently serves 40-minute-old data as if it were live is a correctness bug wearing a reliability costume — and it is exactly how the agent told a customer about a 60-day refund window that had been replaced in July.
Testing the stochastic part
The model's choices are not deterministic, so a test asserting answer == "The refund window is 30 days." will fail on a paraphrase and pass on a lie. Assert invariants and measure rates instead.
| Assert | Do not assert |
|---|---|
| At least one tool call was made | Exactly which tool was chosen first |
| Every factual claim carries a citation | The exact wording of the answer |
| Cited IDs exist in the findings ledger | The number of sources cited |
| The run stayed inside its budgets | The exact iteration count |
| "30 days" appears and "60 days" does not | The sentence containing them |
| Pass rate over N runs is at or above a threshold | That a single run passes |
1@pytest.mark.anyio2@pytest.mark.slow3async def test_refund_question_is_answered_correctly_at_least_9_times_in_10():4 passes = 05 for _ in range(10):6 state = RunState(question=Q)7 answer = await research(Q, mock_manager, state)8 ok = ("30 day" in answer.lower()9 and "60 day" not in answer.lower() # superseded policy10 and "POL-114" in answer11 and all(sid in state.sources() for sid in cited_ids(answer)))12 passes += ok13 assert passes >= 9, f"only {passes}/10 runs were correct"The all(sid in state.sources()) check is the anti-hallucination assertion, and it is the single most valuable line in an agent suite: it fails when the model cites a document ID that no tool call ever returned.
Load testing
Agents fail under concurrency in ways single-run tests cannot reveal: pool exhaustion, upstream rate limits, unbounded queues.
Little's law gives you the sizing before you run anything. Concurrency equals arrival rate times service time. At 20 requests per second with a 250 ms median service time, you need 20 × 0.25 = 5 concurrent connections. If the upstream slows to 2 seconds, the same arrival rate demands 20 × 2 = 40 — and with a pool capped at 10, requests queue, latency climbs, timeouts fire, and the retries make it worse. That collapse is not gradual; it is a cliff.
| Metric | Healthy | Warning | What it means |
|---|---|---|---|
| p50 latency | 120 ms | — | The typical experience |
| p95 latency | 480 ms | 4× p50 or more | Contention is appearing |
| p99 latency | 2,100 ms | 10× p50 or more | Queueing, not slow work |
| Error rate | Below 0.5% | Above 2% | Something is saturated |
| Pool wait time | Near 0 | Above 50 ms | The pool is the bottleneck |
Report percentiles, never averages. A mean of 300 ms is consistent with everyone getting 300 ms and with 95% getting 100 ms while 5% get 4 seconds — and the second case is the one users complain about.
Where suites give false confidence
| Mistake | Why it feels fine | What it hides |
|---|---|---|
| Mocking the MCP client | Fast, deterministic, high coverage | Schema errors, framing bugs, real error shapes |
| Testing only the happy path | The feature works | Every failure branch, which is most of production |
| Asserting exact answer text | Precise-looking | Flaky suite; the team disables it within a month |
| Running each e2e test once | Suite stays fast | A 20% failure rate reads as "passed" |
Mocking sleep away entirely | Tests finish quickly | Whether the backoff schedule is correct at all |
| No test for truncation | Nobody returns 4,300 rows in dev | Context overflow on the day real data arrives |
| Testing with an empty database | Setup is simple | Every bound, cap and pagination path |
If your suite would still pass with every tool returning a hard-coded string, it is testing your mocks, not your agent.
What this means when you build one
The practical shape of a suite that actually protects an MCP agent is lopsided compared with ordinary software, and deliberately so.
Invest most in the deterministic middle. Schema property tests, bound tests, timeout tests, breaker state-machine tests: these are fast, stable, and each one corresponds to a production incident you will otherwise have. They are also the layer where mocks are legitimate, because a fake clock is not a fake system.
Own a mock MCP server, not mock functions. One small server that speaks real MCP and takes a MOCK_FAIL environment variable gives you timeouts, errors, empty results and flakiness on demand, through the real protocol path. It is thirty lines and it replaces most of the mocking in a typical suite.
Run the stochastic tests as a batch with a rate threshold, nightly rather than per-commit, and record the pass rate over time. A drop from 10/10 to 7/10 after a prompt change is a real regression that a single-run test reports as a flake.
Add a test the moment you see a production failure, before fixing it. The 60-day refund answer became one assertion — the answer must not contain "60 day" — and that assertion is now the permanent guard against the class of bug where a stale cache is served as current fact. Suites earn their value from the failures they have already seen, and an agent is a machine for producing new and interesting ones.