Model Context Protocol (MCP)

Build an MCP-Enabled Research Assistant Agent


Three MCP servers, all individually tested and all working: a web search server, an internal document server over pgvector, and a Slack server. Wire them to a model, add a loop, and ask one question — "What changed in our refund policy this quarter, and who discussed it?"

Four minutes later the agent produces 200 words that cite nothing. The transcript shows 63 tool calls. Nineteen of them are calls to a tool named search with near-identical arguments, because two of the three servers exposed a tool called search and the manager's flat dictionary kept whichever registered last. The model asked for internal documents and kept getting the public web, so it rephrased and tried again. And again.

Price it. Forty-seven tools across three servers, at roughly 350 tokens per serialised definition, is 16,450 tokens of tool menu on every single model call. Sixty-three calls, each resending the whole conversation, averages out at something like 30,000 input tokens per call — about 1.9 million input tokens for one question, or roughly 5.70 dollars at an illustrative USD 3 per million input tokens (check your provider's current price list), to produce an answer with no citations.

Every individual server was correct. The system was not. Orchestration is its own engineering problem, and it is where MCP projects actually fail.

The orchestrator's loop across three serversquestionmodelpicks a toolroute tothat serverobservationappendedanswerwith sourcessearch,docs, Slackkeep thesource idState the loop must keep: the transcript, which servers are live, and where every claim came from.
Three working servers do not make an agent — the orchestrator owns routing, state and the stopping rule.

What the orchestrator owns

ResponsibilityWhy it cannot live in a serverWhat happens if nobody owns it
Namespacing tool namesA server cannot know what other servers existCollisions route calls to the wrong system
Routing calls to serversOnly the host holds all the connectionsCalls fail or go to the wrong server
Budgets: iterations, tokens, wall clockA server sees one call, never the loopRunaway loops and surprise invoices
Deduplicating repeated callsServers are stateless across callsThe same query is run nineteen times
Accumulating findings and citationsServers return fragments, not a reportAnswers with no provenance
Degrading when a server is downA dead server cannot report itselfOne failed connection kills the whole run

An agent is not a model plus tools. It is a model plus tools plus a budget, a router, and a memory — and the last three are yours to write.

The client manager

The manager connects to each configured server, prefixes every tool name with the server's alias, and keeps the map from prefixed name back to the client for that server. It uses the Client class from version 2 of the MCP Python SDK, which works out on its own whether a server speaks the stateless 2026-07-28 protocol or an older handshake-based one.

Python
import asyncio, json, hashlib, os, sys, timefrom contextlib import AsyncExitStackfrom dataclasses import dataclass, fieldfrom mcp import Client, StdioServerParameters@dataclassclass ServerSpec:    alias: str                    # short, becomes the tool-name prefix    command: str    args: list[str]    env: dict[str, str] = field(default_factory=dict)    required: bool = False        # if True, startup fails when it failsclass MCPManager:    def __init__(self):        self.clients: dict[str, Client] = {}        self.tool_owner: dict[str, str] = {}      # "docs__search" -> "docs"        self.tool_defs: list[dict] = []           # what the model is shown        self.unavailable: dict[str, str] = {}     # alias -> reason        self._stack = AsyncExitStack()    async def connect(self, spec: ServerSpec) -> None:        params = StdioServerParameters(command=spec.command, args=spec.args,                                       env=spec.env or None)        client = await self._stack.enter_async_context(Client(params))        listing = await asyncio.wait_for(client.list_tools(), timeout=20)        for tool in listing.tools:            qualified = f"{spec.alias}__{tool.name}"     # namespace, always            if qualified in self.tool_owner:                raise RuntimeError(f"duplicate tool name {qualified}")            self.tool_owner[qualified] = spec.alias            self.tool_defs.append({                "name": qualified,                "description": tool.description or "",                "input_schema": tool.input_schema,            })        self.clients[spec.alias] = client    async def start(self, specs: list[ServerSpec]) -> None:        await self._stack.__aenter__()        for spec in specs:            try:                await self.connect(spec)            except Exception as e:                if spec.required:                    raise                self.unavailable[spec.alias] = str(e)   # degrade, do not die                print(f"server {spec.alias} unavailable: {e}", file=sys.stderr)    async def call(self, qualified: str, arguments: dict, timeout: float = 45.0) -> str:        alias = self.tool_owner.get(qualified)        if alias is None:            available = ", ".join(sorted(self.tool_owner)[:10])            return f"No such tool: {qualified}. Available tools include: {available}"        client = self.clients[alias]        bare = qualified.split("__", 1)[1]        try:            result = await asyncio.wait_for(                client.call_tool(bare, arguments), timeout=timeout)        except asyncio.TimeoutError:            return (f"{qualified} did not respond within {timeout:.0f}s. "                    f"Try a narrower query or a different source.")        except Exception as e:            return f"{qualified} failed: {type(e).__name__}: {e}"        text = "\n".join(c.text for c in result.content if c.type == "text")        return text or "(the tool returned no content)"

Four decisions there are worth defending. The double-underscore prefix keeps names inside the character set tool APIs accept while staying readable to the model — docs__search versus web__search is unambiguous in a way that search and search_2 never are. The duplicate check raises at startup, so a collision is a deployment failure rather than a silent misroute at 3am. The non-required servers degrade, so Slack being down costs you Slack rather than the whole assistant. And every error path returns a sentence instead of raising, because the model is the thing that has to decide what to do next, and it can only decide from text it can read.

Wiring the three servers

Python
SPECS = [    ServerSpec(alias="web",   command="uv",               args=["run", "--directory", "/srv/web-search", "python", "server.py"],               env={"GOOGLE_API_KEY": os.environ["GOOGLE_API_KEY"]}),    ServerSpec(alias="docs",  command="uv",               args=["run", "--directory", "/srv/doc-store", "python", "server.py"],               env={"DATABASE_URL": os.environ["DOCS_DSN"]},               required=True),                       # without documents there is no product    ServerSpec(alias="slack", command="uv",               args=["run", "--directory", "/srv/slack-mcp", "python", "server.py"],               env={"SLACK_BOT_TOKEN": os.environ["SLACK_BOT_TOKEN"]}),]

Before handing that tool list to the model, count it. Three servers exposing 47 tools costs about 16,450 tokens per call. Cut to the 12 tools a research assistant genuinely uses and the menu costs about 4,200 — a saving of roughly 12,000 tokens on every model call in the run. Over nine calls that is 110,000 tokens saved, and selection accuracy goes up because the model is choosing between twelve distinguishable options rather than forty-seven overlapping ones.

Python
ALLOWED = {    "web__web_search", "web__fetch_page",    "docs__semantic_search", "docs__get_document", "docs__list_collections",    "slack__search_messages", "slack__get_thread", "slack__list_channels",}manager.tool_defs = [t for t in manager.tool_defs if t["name"] in ALLOWED]

The research loop

The loop is small. Its value is entirely in the guards.

Python
import anthropicclient = anthropic.AsyncAnthropic()MODEL = os.environ.get("RESEARCH_MODEL", "claude-sonnet-5")   # a good default; set per deploymentSYSTEM = """You are a research assistant with access to three sources.- docs__semantic_search: internal company documents. Authoritative for policy.- web__web_search: the public web. Use for external context only.- slack__search_messages: internal discussion. Good for who said what and when.Rules:- Answer only from tool results. If the tools do not support a claim, say so.- Cite every factual claim with the document ID, URL or Slack permalink.- Prefer internal documents over the web for anything about company policy.- Run independent searches in parallel in a single turn where you can.- Stop searching once you can answer; do not gather more for its own sake."""@dataclassclass Budget:    max_iterations: int = 12    max_seconds: float = 180.0    max_input_tokens: int = 400_000    started: float = field(default_factory=time.monotonic)    used_input: int = 0    iterations: int = 0    def exhausted(self) -> str | None:        if self.iterations >= self.max_iterations:            return f"iteration limit ({self.max_iterations})"        if time.monotonic() - self.started > self.max_seconds:            return f"time limit ({self.max_seconds:.0f}s)"        if self.used_input > self.max_input_tokens:            return f"token budget ({self.max_input_tokens:,})"        return Noneasync def research(question: str, manager: MCPManager, state: "RunState") -> str:    messages = [{"role": "user", "content": question}]    budget = Budget()    while True:        reason = budget.exhausted()        if reason:            messages.append({"role": "user", "content":                f"Budget reached ({reason}). Write the best answer you can from "                f"what you already have, and state clearly what remains unverified."})        resp = await client.messages.create(            model=MODEL,            max_tokens=4096,            system=SYSTEM,            tools=manager.tool_defs,            messages=messages,        )        budget.used_input += resp.usage.input_tokens        budget.iterations += 1        messages.append({"role": "assistant", "content": resp.content})        if resp.stop_reason != "tool_use" or reason:            return "".join(b.text for b in resp.content if b.type == "text")        calls = [b for b in resp.content if b.type == "tool_use"]        results = await asyncio.gather(*[            execute_one(manager, state, c) for c in calls])        messages.append({"role": "user", "content": results})async def execute_one(manager: MCPManager, state: "RunState", block) -> dict:    key = hashlib.sha256(        f"{block.name}:{json.dumps(block.input, sort_keys=True)}".encode()).hexdigest()    if key in state.seen_calls:        text = ("You already ran this exact call and received the result above. "                "Change the query, try a different source, or answer with what you have.")    else:        state.seen_calls.add(key)        started = time.monotonic()        text = await manager.call(block.name, block.input)        state.record(block.name, block.input, text, time.monotonic() - started)    return {"type": "tool_result", "tool_use_id": block.id,            "content": text[:20_000], "is_error": text.startswith(("No such tool",))}

The duplicate guard deserves attention, because it is the specific fix for the 19-identical-searches failure. Hashing the tool name plus its canonicalised arguments catches exact repeats, and instead of silently returning the cached result it tells the model that it is repeating itself. Models loop when they cannot tell that an approach has already failed; naming the loop breaks it.

Parallel execution matters too. Three independent searches issued in one turn run concurrently, so the wall-clock cost is the slowest of the three rather than their sum — 1.4 seconds instead of 4.1 — and, more importantly, they cost one model call instead of three.

State the loop must keep

Python
@dataclassclass Finding:    tool: str    query: str    source_id: str    excerpt: str    seconds: float@dataclassclass RunState:    question: str    seen_calls: set[str] = field(default_factory=set)    findings: list[Finding] = field(default_factory=list)    errors: list[tuple[str, str]] = field(default_factory=list)    def record(self, tool: str, args: dict, text: str, seconds: float) -> None:        if text.startswith("No such tool") or any(                marker in text[:80] for marker in ("did not respond", "failed:")):            self.errors.append((tool, text[:200]))            return        for source_id, excerpt in extract_citations(text):            self.findings.append(Finding(tool, json.dumps(args)[:200],                                         source_id, excerpt[:400], seconds))    def sources(self) -> list[str]:        seen, out = set(), []        for f in self.findings:                 # preserve discovery order            if f.source_id not in seen:                seen.add(f.source_id)                out.append(f.source_id)        return out

Keeping findings outside the message list is not redundancy. The message list is what the model sees and it will be truncated, summarised or dropped when it grows; the findings ledger is what you use to verify citations, render a sources section, and detect the failure where the agent produced a confident answer from zero successful tool calls. That check is one line and it catches the worst class of bug in the whole system:

Python
if not state.findings and state.errors:    return ("I could not retrieve any sources for this question. "            f"{len(state.errors)} tool call(s) failed: "            + "; ".join(f"{t}: {msg}" for t, msg in state.errors[:3]))

A mock server for development

Developing against live servers is slow, costs quota, and cannot reproduce failures on demand. A mock server that speaks real MCP over stdio fixes all three, and it takes about thirty lines.

Python
from mcp.server.mcpserver import MCPServerimport os, asyncio, randommcp = MCPServer("mock-docs")FAIL_MODE = os.environ.get("MOCK_FAIL", "none")   # none|timeout|error|empty|flakyCORPUS = {    "POL-114": "Refund policy v4, effective 2026-07-01. Refunds are available "               "within 30 days of delivery, reduced from 60 days in v3.",    "POL-098": "Refund policy v3, effective 2025-01-01. Refunds within 60 days.",    "ENG-421": "Runbook: refund reconciliation job runs nightly at 02:00 UTC.",}@mcp.tool()async def semantic_search(query: str, k: int = 3) -> str:    """Search internal documents. Returns document IDs and excerpts."""    if FAIL_MODE == "timeout":        await asyncio.sleep(120)    if FAIL_MODE == "error":        raise RuntimeError("simulated backend failure")    if FAIL_MODE == "empty":        return "No documents matched that query."    if FAIL_MODE == "flaky" and random.random() < 0.4:        raise RuntimeError("simulated transient failure")    hits = [(doc_id, text) for doc_id, text in CORPUS.items()            if any(w in text.lower() for w in query.lower().split())][:k]    if not hits:        return "No documents matched that query."    return "\n\n".join(f"[{doc_id}] {text}" for doc_id, text in hits)if __name__ == "__main__":    mcp.run()

Because it is a real MCP server, swapping it in is a one-line config change and the whole path — version detection, tool listing, JSON-RPC framing, argument validation — is exercised exactly as in production. That is what distinguishes a mock server from a mocked client object: patching manager.call in a test proves nothing about whether your schemas are valid.

Two traces

The happy path

StepWhat happensElapsedInput tokens
1Model receives the question and 12 tool definitions0.0 s4,900
2Emits three tool calls in one turn: docs, slack, web2.1 s—
3All three run concurrently; slowest is web at 1.4 s3.5 s—
4Model reads 3 results (about 3,100 tokens total)3.5 s8,000
5Sees POL-114 supersedes POL-098; fetches the Slack thread5.2 s—
6Thread returns 14 messages, about 900 tokens6.0 s—
7Writes the answer with four citations9.4 s10,300

Three model calls, four tool calls, about 23,200 input tokens in total — roughly 0.07 dollars at the same illustrative price, against the 5.70 dollars the unguarded version spent. The saving comes from three things and none of them is a better model: a smaller tool menu, parallel calls collapsing three round trips into one, and a loop that stops when it has enough.

Error recovery

The same question, but the Slack server's token has expired.

Text
iter 1  model -> docs__semantic_search("refund policy change 2026")                slack__search_messages("refund policy")        docs  -> [POL-114] Refund policy v4 ... 30 days, reduced from 60.        slack -> slack__search_messages failed: The Slack token is invalid or                 revoked. Reconnect the workspace.iter 2  model -> docs__semantic_search("refund policy discussion approval")        docs  -> [POL-114] ... (same document, no discussion thread)iter 3  model -> final answer:        "The refund window changed from 60 days to 30 days, effective         2026-07-01 (source: POL-114, superseding POL-098). I could not         determine who discussed the change: the Slack integration is         disconnected and needs to be reauthorised."

That is the behaviour you are engineering for. The agent lost one source, tried a reasonable alternative, and then said what it could not determine and why rather than filling the gap. Two things made it possible: the manager returned the vendor's real reason as readable text instead of raising, and the system prompt said "if the tools do not support a claim, say so". Remove either and the same run produces a confident, invented answer about who approved the policy.

Where these agents go wrong

FailureWhat you seeRoot causeFix
Tool-name collisionResults from the wrong systemFlat name map across serversPrefix by alias; raise on duplicates
Infinite loopSame call repeated with tiny variationsNo memory of attemptsHash calls; tell the model it repeated
Context overflow mid-runFailure at iteration 8, never iteration 1Unbounded tool resultsTruncate results; cap tool count; budget tokens
Confident answer, zero sourcesFluent text, no citationsErrors swallowed into empty resultsCheck the findings ledger before returning
One dead server kills the runStartup crashAll servers treated as requiredMark most as optional; tell the model what is missing
Serial when it could be parallelFour-minute latency, low token useOne call per turnPrompt for parallel calls; gather them
Stale cached tool listCalls a tool the server removedIgnoring list_changedRe-list on the notification

Every guard in an agent loop exists because a model, given no reason to stop, will not stop. Budgets are not pessimism; they are the control system.

What this means when you build one

The instinct when an agent underperforms is to improve the prompt. In practice the highest-leverage changes are almost always structural, and they are boring.

Curate the tool list ruthlessly. Twelve well-described tools beat forty-seven, on both cost and accuracy, and the curation is a config file rather than a research problem. If two tools would be chosen by the same sentence from a user, merge them or delete one.

Make the loop observable before you make it clever. Log every iteration with the tool name, argument hash, duration, result size and running token count. The 63-call disaster was diagnosable in ninety seconds once those lines existed and essentially undiagnosable before.

Treat every tool result as untrusted text that will be read by a model. Truncate it, and remember that a document fetched from the web can contain instructions aimed at your agent. The tool-result path is the main injection surface in a system like this, which is a strong argument for keeping high-privilege write tools out of any loop that also reads the open web.

Decide what "done" means, explicitly. The loop above stops on three budgets and on the model declining to call a tool. That is a policy, and writing it down as one is the difference between an agent that answers in nine seconds and one that researches until someone kills the process.