Model Context Protocol (MCP)

Testing and Reliability for MCP Agents


A team has 340 tests for their MCP agent. All green, on every commit, for six weeks. In production the agent fails roughly one run in six: sometimes it hangs for two minutes and returns nothing, sometimes it answers confidently with no sources, and once it told a customer their refund window was 60 days when the policy had said 30 days since July.

Read the suite and the reason is immediate. Every test patches the client manager:

Python
manager.call = AsyncMock(return_value="[POL-114] Refunds within 30 days.")

That line asserts that if a tool returns a well-formed string, the agent handles it. Not one test ever sent a real JSON-RPC request, sent an argument through a real schema validator, saw an HTTP 429, waited out a timeout, or watched what happens when the tool returns 40,000 tokens. The suite tests the glue between mocks. The production failures all live in the parts that were mocked away.

Testing an agent is not harder than testing ordinary software, but the boundaries are in different places, and putting them in the usual places produces exactly this: high coverage and low confidence.

Where the test pyramid bends for agentsUnit: handlersand their boundsCataloguechecked as dataIntegration onreal transportTimeout, retry,breaker, fallbackStochasticruns, scoredtopbottom340 green tests missed a two-minute hang because nothing exercised the transport under failure.
Most agent failures live above the unit layer, which is exactly where the cheap tests stop.

Why the pyramid bends

The critical observation is that an agent system has one stochastic component and many deterministic ones, and they need completely different treatment.

LayerDeterministic?How to test itAssert on
Server handlersYesUnit tests, direct callsExact outputs
Tool schemasYesProperty tests over the catalogueStructural rules
Transport and discoveryYesIntegration, real subprocessServer answers, tools list
Retry / breaker / timeoutYes, with a fake clockFault injectionNumber and timing of attempts
Model tool selectionNoRepeated runs against fixturesRates and invariants, never exact text
Final answer wordingNoProperty checks, spot reviewCitations present, no unsupported claims

Assert exact behaviour everywhere the system is deterministic, and assert only properties and rates where it is not. Mixing those up gives you either a flaky suite or a useless one.

Unit-testing the server

The MCP Python SDK can run a client and server in the same process: pass the server object itself to Client and it connects over an in-memory transport. This is the sweet spot: no subprocess, no network, but real protocol.

Python
import pytestfrom mcp import Clientfrom server import mcp          # the MCPServer instance under test@pytest.mark.anyioasync def test_tool_is_advertised_with_a_usable_schema():    async with Client(mcp) as client:        tools = {t.name: t for t in (await client.list_tools()).tools}        assert "search_orders" in tools        schema = tools["search_orders"].input_schema        assert schema["required"] == ["customer_email"]        assert schema["properties"]["limit"]["maximum"] == 50        assert "shipped" in schema["properties"]["status"]["anyOf"][0]["enum"]@pytest.mark.anyioasync def test_bad_date_returns_a_correctable_message_not_an_exception():    async with Client(mcp) as client:        result = await client.call_tool(            "search_orders", {"customer_email": "a@b.com", "since": "last week"})        text = result.content[0].text        assert "ISO-8601" in text          # tells the model the format        assert "2026-" in text             # and shows an example

That second test encodes a rule most suites miss: an invalid argument must produce a message the model can act on. Asserting only that the call did not crash lets someone "fix" a bug by returning the string "error", which passes the test and destroys the agent's ability to recover.

Bounds are the tests people skip

Python
@pytest.mark.anyioasync def test_result_is_bounded_even_when_the_data_is_not(seeded_db_with_4300_orders):    async with Client(mcp) as client:        result = await client.call_tool(            "search_orders", {"customer_email": "bulk@example.com", "limit": 50})        text = result.content[0].text        assert len(text) <= 40_000        assert text.count("\n") <= 55          # 50 rows plus header and notice        assert "more matched" in text          # the model is told it was truncated

A truncation that does not announce itself is worse than no truncation, because the model treats a partial result as complete and reasons from it.

Testing the catalogue as data

Some of the highest-value tests iterate over every tool and check structural rules. They cost twenty lines and catch a whole class of mistakes at the moment they are introduced.

Python
import refrom jsonschema import Draft202012ValidatorNAME_RE = re.compile(r"^[a-z][a-z0-9_]{2,63}$")@pytest.mark.anyioasync def test_every_tool_definition_is_well_formed():    async with Client(mcp) as client:        tools = (await client.list_tools()).tools    assert len({t.name for t in tools}) == len(tools), "duplicate tool names"    for t in tools:        assert NAME_RE.match(t.name), f"bad tool name: {t.name}"        Draft202012Validator.check_schema(t.input_schema)  # the schema itself is valid        desc = (t.description or "").strip()        assert len(desc) >= 40, f"{t.name}: description too thin to choose from"        assert len(desc) <= 1024, f"{t.name}: description will bloat every prompt"        props = t.input_schema.get("properties", {})        for field, spec in props.items():            assert spec.get("description"), f"{t.name}.{field} has no description"            if spec.get("type") == "integer":                assert "maximum" in spec, f"{t.name}.{field} is an unbounded integer"    total = sum(len(t.name) + len(t.description or "") +                len(str(t.input_schema)) for t in tools) // 4    assert total < 6_000, f"tool catalogue is ~{total} tokens per model call"

That last assertion is a budget guard in test form. A catalogue that quietly grows from 4,200 to 16,000 tokens raises the cost of every model call in the system, and nothing else in a normal CI pipeline notices.

Integration: use the real transport

In-memory connections skip subprocess launch, environment handling and stream framing — which is where the "works on my machine" failures live. At least one test must spawn the server for real.

Python
from mcp import Client, StdioServerParameters@pytest.mark.anyioasync def test_server_starts_from_a_clean_environment_and_lists_tools():    params = StdioServerParameters(        command="uv", args=["run", "python", "server.py"],        env={"DATABASE_URL": TEST_DSN, "PATH": os.environ["PATH"]},   # nothing else    )    async with Client(params) as client:        listing = await asyncio.wait_for(client.list_tools(), timeout=10)        assert client.protocol_version == "2026-07-28"   # not a legacy fallback        assert client.server_capabilities.tools is not None        assert {t.name for t in listing.tools} == EXPECTED_TOOL_NAMES

Passing a minimal env is the point of the test. The protocol-version assertion is a cheap guard too: if a dependency pin drags the server back to an SDK that only speaks the old handshake, the client silently falls back, and this line tells you. Servers routinely work locally because the developer's shell exports a variable that the deployment does not, and this test fails loudly on exactly that.

The four failure patterns

Real MCP agents fail in four recognisable ways. Each needs its own mechanism and its own test.

Timeouts

Every outbound call needs a deadline, and the deadline must be tested with a server that genuinely hangs — not with a mock that raises TimeoutError, which proves only that your except clause has correct syntax.

Python
@pytest.mark.anyioasync def test_hanging_server_is_abandoned_with_a_useful_message():    manager = await start_manager(MOCK_FAIL="timeout")     # the mock sleeps 120 s    started = time.monotonic()    text = await manager.call("docs__semantic_search", {"query": "refunds"},                              timeout=2.0)    elapsed = time.monotonic() - started    assert 2.0 <= elapsed < 3.0, "timeout was not enforced"    assert "did not respond" in text    assert "narrower query" in text                         # actionable for the model

Timeout budgets must nest. If the agent's overall budget is 180 seconds and a single tool may take 45, then four sequential timeouts exhaust the run before the model writes anything. Set per-call timeouts so that the worst realistic sequence still leaves time to answer: with a 12-iteration cap, 45 seconds each is 540 seconds of potential blocking against a 180-second budget — the numbers do not fit, and 15 seconds per call is the honest ceiling.

Retry with backoff

Test the schedule, not just the outcome. Capture the sleeps.

Python
@pytest.mark.anyioasync def test_backoff_grows_and_is_jittered(monkeypatch):    waits: list[float] = []    async def fake_sleep(s): waits.append(s)    monkeypatch.setattr(asyncio, "sleep", fake_sleep)    attempts = 0    async def flaky():        nonlocal attempts        attempts += 1        if attempts < 4:            raise Retryable("429")        return "ok"    assert await call_with_retry(flaky) == "ok"    assert attempts == 4    assert len(waits) == 3    # Each wait is bounded by the exponential ceiling for its attempt...    assert waits[0] <= 1.0 and waits[1] <= 2.0 and waits[2] <= 4.0    # ...and jitter means they are not all identical.    assert len(set(waits)) > 1

The final assertion is the one that catches the outage from the opening: without jitter every client waits the identical interval, and the retries arrive as a synchronised burst that re-triggers the throttle.

Circuit breaker

Retries handle a blip. They are the wrong tool for an upstream that is down, because every call still pays the full timeout before failing. Work the numbers: an upstream out for ten minutes, an agent making 20 calls a minute, a 45-second timeout on each — that is 200 calls each blocking a worker for 45 seconds, or 2.5 hours of worker time spent waiting for a service you already know is broken.

A circuit breaker converts that into 5 slow failures followed by instant ones.

Python
import timeclass CircuitBreaker:    """CLOSED -> (5 failures in 60 s) -> OPEN -> (after cooldown) -> HALF_OPEN"""    def __init__(self, threshold=5, window=60.0, cooldown=30.0, max_cooldown=300.0):        self.threshold, self.window = threshold, window        self.cooldown, self.max_cooldown = cooldown, max_cooldown        self.state, self.failures, self.opened_at = "CLOSED", [], 0.0        self.current_cooldown = cooldown    def allow(self) -> bool:        now = time.monotonic()        if self.state == "OPEN":            if now - self.opened_at >= self.current_cooldown:                self.state = "HALF_OPEN"      # let exactly one trial through                return True            return False        return True    def record(self, ok: bool) -> None:        now = time.monotonic()        if ok:            self.state, self.failures = "CLOSED", []            self.current_cooldown = self.cooldown        # reset the escalation            return        if self.state == "HALF_OPEN":                    # trial failed: back off harder            self.state, self.opened_at = "OPEN", now            self.current_cooldown = min(self.current_cooldown * 2, self.max_cooldown)            return        self.failures = [t for t in self.failures if now - t < self.window] + [now]        if len(self.failures) >= self.threshold:            self.state, self.opened_at = "OPEN", now
Python
@pytest.mark.anyioasync def test_breaker_opens_then_probes_then_closes(monkeypatch):    clock = [1000.0]    monkeypatch.setattr(time, "monotonic", lambda: clock[0])    cb = CircuitBreaker(threshold=5, window=60.0, cooldown=30.0)    for _ in range(5):        assert cb.allow()        cb.record(ok=False)    assert cb.state == "OPEN"    assert cb.allow() is False               # fails instantly, no timeout paid    clock[0] += 31.0    assert cb.allow() is True                # one probe permitted    assert cb.state == "HALF_OPEN"    cb.record(ok=False)    assert cb.state == "OPEN"    assert cb.current_cooldown == 60.0       # escalated 30 -> 60    clock[0] += 61.0    assert cb.allow() is True    cb.record(ok=True)    assert cb.state == "CLOSED"

Fallback

When the breaker is open, something still has to answer. A fallback chain degrades in a defined order and — critically — tells the caller which tier answered.

Python
async def search_with_fallback(query: str) -> tuple[str, str]:    """Returns (text, tier). The tier travels with the result so the model    can qualify its answer instead of presenting stale data as current."""    if breakers["docs"].allow():        try:            text = await manager.call("docs__semantic_search", {"query": query})            breakers["docs"].record(ok=True)            return text, "primary"        except Exception:            breakers["docs"].record(ok=False)    cached = cache.get(query)                       # tier 2: possibly stale    if cached:        return f"{cached}\n\n(from cache, {cache.age(query)//60} minutes old)", "cache"    if breakers["web"].allow():                     # tier 3: different source        try:            return await manager.call("web__web_search", {"query": query}), "degraded"        except Exception:            breakers["web"].record(ok=False)    return ("No source is currently reachable for this query. "            "State that the information could not be verified."), "unavailable"

Test that each tier is reachable and that the tier label is truthful. A fallback that silently serves 40-minute-old data as if it were live is a correctness bug wearing a reliability costume — and it is exactly how the agent told a customer about a 60-day refund window that had been replaced in July.

Testing the stochastic part

The model's choices are not deterministic, so a test asserting answer == "The refund window is 30 days." will fail on a paraphrase and pass on a lie. Assert invariants and measure rates instead.

AssertDo not assert
At least one tool call was madeExactly which tool was chosen first
Every factual claim carries a citationThe exact wording of the answer
Cited IDs exist in the findings ledgerThe number of sources cited
The run stayed inside its budgetsThe exact iteration count
"30 days" appears and "60 days" does notThe sentence containing them
Pass rate over N runs is at or above a thresholdThat a single run passes
Python
@pytest.mark.anyio@pytest.mark.slowasync def test_refund_question_is_answered_correctly_at_least_9_times_in_10():    passes = 0    for _ in range(10):        state = RunState(question=Q)        answer = await research(Q, mock_manager, state)        ok = ("30 day" in answer.lower()              and "60 day" not in answer.lower()          # superseded policy              and "POL-114" in answer              and all(sid in state.sources() for sid in cited_ids(answer)))        passes += ok    assert passes >= 9, f"only {passes}/10 runs were correct"

The all(sid in state.sources()) check is the anti-hallucination assertion, and it is the single most valuable line in an agent suite: it fails when the model cites a document ID that no tool call ever returned.

Load testing

Agents fail under concurrency in ways single-run tests cannot reveal: pool exhaustion, upstream rate limits, unbounded queues.

Little's law gives you the sizing before you run anything. Concurrency equals arrival rate times service time. At 20 requests per second with a 250 ms median service time, you need 20 × 0.25 = 5 concurrent connections. If the upstream slows to 2 seconds, the same arrival rate demands 20 × 2 = 40 — and with a pool capped at 10, requests queue, latency climbs, timeouts fire, and the retries make it worse. That collapse is not gradual; it is a cliff.

MetricHealthyWarningWhat it means
p50 latency120 ms—The typical experience
p95 latency480 ms4× p50 or moreContention is appearing
p99 latency2,100 ms10× p50 or moreQueueing, not slow work
Error rateBelow 0.5%Above 2%Something is saturated
Pool wait timeNear 0Above 50 msThe pool is the bottleneck

Report percentiles, never averages. A mean of 300 ms is consistent with everyone getting 300 ms and with 95% getting 100 ms while 5% get 4 seconds — and the second case is the one users complain about.

Where suites give false confidence

MistakeWhy it feels fineWhat it hides
Mocking the MCP clientFast, deterministic, high coverageSchema errors, framing bugs, real error shapes
Testing only the happy pathThe feature worksEvery failure branch, which is most of production
Asserting exact answer textPrecise-lookingFlaky suite; the team disables it within a month
Running each e2e test onceSuite stays fastA 20% failure rate reads as "passed"
Mocking sleep away entirelyTests finish quicklyWhether the backoff schedule is correct at all
No test for truncationNobody returns 4,300 rows in devContext overflow on the day real data arrives
Testing with an empty databaseSetup is simpleEvery bound, cap and pagination path

If your suite would still pass with every tool returning a hard-coded string, it is testing your mocks, not your agent.

What this means when you build one

The practical shape of a suite that actually protects an MCP agent is lopsided compared with ordinary software, and deliberately so.

Invest most in the deterministic middle. Schema property tests, bound tests, timeout tests, breaker state-machine tests: these are fast, stable, and each one corresponds to a production incident you will otherwise have. They are also the layer where mocks are legitimate, because a fake clock is not a fake system.

Own a mock MCP server, not mock functions. One small server that speaks real MCP and takes a MOCK_FAIL environment variable gives you timeouts, errors, empty results and flakiness on demand, through the real protocol path. It is thirty lines and it replaces most of the mocking in a typical suite.

Run the stochastic tests as a batch with a rate threshold, nightly rather than per-commit, and record the pass rate over time. A drop from 10/10 to 7/10 after a prompt change is a real regression that a single-run test reports as a flake.

Add a test the moment you see a production failure, before fixing it. The 60-day refund answer became one assertion — the answer must not contain "60 day" — and that assertion is now the permanent guard against the class of bug where a stale cache is served as current fact. Suites earn their value from the failures they have already seen, and an agent is a machine for producing new and interesting ones.