Course Content
Model Context Protocol (MCP)
3 sections · 8 lessons
Build an MCP-Enabled Research Assistant Agent
Three MCP servers, all individually tested and all working: a web search server, an internal document server over pgvector, and a Slack server. Wire them to a model, add a loop, and ask one question — "What changed in our refund policy this quarter, and who discussed it?"
Four minutes later the agent produces 200 words that cite nothing. The transcript shows 63 tool calls. Nineteen of them are calls to a tool named search with near-identical arguments, because two of the three servers exposed a tool called search and the manager's flat dictionary kept whichever registered last. The model asked for internal documents and kept getting the public web, so it rephrased and tried again. And again.
Price it. Forty-seven tools across three servers, at roughly 350 tokens per serialised definition, is 16,450 tokens of tool menu on every single model call. Sixty-three calls, each resending the whole conversation, averages out at something like 30,000 input tokens per call — about 1.9 million input tokens for one question, or roughly 5.70 dollars at an illustrative USD 3 per million input tokens (check your provider's current price list), to produce an answer with no citations.
Every individual server was correct. The system was not. Orchestration is its own engineering problem, and it is where MCP projects actually fail.
What the orchestrator owns
| Responsibility | Why it cannot live in a server | What happens if nobody owns it |
|---|---|---|
| Namespacing tool names | A server cannot know what other servers exist | Collisions route calls to the wrong system |
| Routing calls to servers | Only the host holds all the connections | Calls fail or go to the wrong server |
| Budgets: iterations, tokens, wall clock | A server sees one call, never the loop | Runaway loops and surprise invoices |
| Deduplicating repeated calls | Servers are stateless across calls | The same query is run nineteen times |
| Accumulating findings and citations | Servers return fragments, not a report | Answers with no provenance |
| Degrading when a server is down | A dead server cannot report itself | One failed connection kills the whole run |
An agent is not a model plus tools. It is a model plus tools plus a budget, a router, and a memory — and the last three are yours to write.
The client manager
The manager connects to each configured server, prefixes every tool name with the server's alias, and keeps the map from prefixed name back to the client for that server. It uses the Client class from version 2 of the MCP Python SDK, which works out on its own whether a server speaks the stateless 2026-07-28 protocol or an older handshake-based one.
1import asyncio, json, hashlib, os, sys, time2from contextlib import AsyncExitStack3from dataclasses import dataclass, field4from mcp import Client, StdioServerParameters56@dataclass7class ServerSpec:8 alias: str # short, becomes the tool-name prefix9 command: str10 args: list[str]11 env: dict[str, str] = field(default_factory=dict)12 required: bool = False # if True, startup fails when it fails1314class MCPManager:15 def __init__(self):16 self.clients: dict[str, Client] = {}17 self.tool_owner: dict[str, str] = {} # "docs__search" -> "docs"18 self.tool_defs: list[dict] = [] # what the model is shown19 self.unavailable: dict[str, str] = {} # alias -> reason20 self._stack = AsyncExitStack()2122 async def connect(self, spec: ServerSpec) -> None:23 params = StdioServerParameters(command=spec.command, args=spec.args,24 env=spec.env or None)25 client = await self._stack.enter_async_context(Client(params))26 listing = await asyncio.wait_for(client.list_tools(), timeout=20)27 for tool in listing.tools:28 qualified = f"{spec.alias}__{tool.name}" # namespace, always29 if qualified in self.tool_owner:30 raise RuntimeError(f"duplicate tool name {qualified}")31 self.tool_owner[qualified] = spec.alias32 self.tool_defs.append({33 "name": qualified,34 "description": tool.description or "",35 "input_schema": tool.input_schema,36 })37 self.clients[spec.alias] = client3839 async def start(self, specs: list[ServerSpec]) -> None:40 await self._stack.__aenter__()41 for spec in specs:42 try:43 await self.connect(spec)44 except Exception as e:45 if spec.required:46 raise47 self.unavailable[spec.alias] = str(e) # degrade, do not die48 print(f"server {spec.alias} unavailable: {e}", file=sys.stderr)4950 async def call(self, qualified: str, arguments: dict, timeout: float = 45.0) -> str:51 alias = self.tool_owner.get(qualified)52 if alias is None:53 available = ", ".join(sorted(self.tool_owner)[:10])54 return f"No such tool: {qualified}. Available tools include: {available}"55 client = self.clients[alias]56 bare = qualified.split("__", 1)[1]57 try:58 result = await asyncio.wait_for(59 client.call_tool(bare, arguments), timeout=timeout)60 except asyncio.TimeoutError:61 return (f"{qualified} did not respond within {timeout:.0f}s. "62 f"Try a narrower query or a different source.")63 except Exception as e:64 return f"{qualified} failed: {type(e).__name__}: {e}"65 text = "\n".join(c.text for c in result.content if c.type == "text")66 return text or "(the tool returned no content)"Four decisions there are worth defending. The double-underscore prefix keeps names inside the character set tool APIs accept while staying readable to the model — docs__search versus web__search is unambiguous in a way that search and search_2 never are. The duplicate check raises at startup, so a collision is a deployment failure rather than a silent misroute at 3am. The non-required servers degrade, so Slack being down costs you Slack rather than the whole assistant. And every error path returns a sentence instead of raising, because the model is the thing that has to decide what to do next, and it can only decide from text it can read.
Wiring the three servers
1SPECS = [2 ServerSpec(alias="web", command="uv",3 args=["run", "--directory", "/srv/web-search", "python", "server.py"],4 env={"GOOGLE_API_KEY": os.environ["GOOGLE_API_KEY"]}),5 ServerSpec(alias="docs", command="uv",6 args=["run", "--directory", "/srv/doc-store", "python", "server.py"],7 env={"DATABASE_URL": os.environ["DOCS_DSN"]},8 required=True), # without documents there is no product9 ServerSpec(alias="slack", command="uv",10 args=["run", "--directory", "/srv/slack-mcp", "python", "server.py"],11 env={"SLACK_BOT_TOKEN": os.environ["SLACK_BOT_TOKEN"]}),12]Before handing that tool list to the model, count it. Three servers exposing 47 tools costs about 16,450 tokens per call. Cut to the 12 tools a research assistant genuinely uses and the menu costs about 4,200 — a saving of roughly 12,000 tokens on every model call in the run. Over nine calls that is 110,000 tokens saved, and selection accuracy goes up because the model is choosing between twelve distinguishable options rather than forty-seven overlapping ones.
1ALLOWED = {2 "web__web_search", "web__fetch_page",3 "docs__semantic_search", "docs__get_document", "docs__list_collections",4 "slack__search_messages", "slack__get_thread", "slack__list_channels",5}6manager.tool_defs = [t for t in manager.tool_defs if t["name"] in ALLOWED]The research loop
The loop is small. Its value is entirely in the guards.
1import anthropic23client = anthropic.AsyncAnthropic()4MODEL = os.environ.get("RESEARCH_MODEL", "claude-sonnet-5") # a good default; set per deployment56SYSTEM = """You are a research assistant with access to three sources.78- docs__semantic_search: internal company documents. Authoritative for policy.9- web__web_search: the public web. Use for external context only.10- slack__search_messages: internal discussion. Good for who said what and when.1112Rules:13- Answer only from tool results. If the tools do not support a claim, say so.14- Cite every factual claim with the document ID, URL or Slack permalink.15- Prefer internal documents over the web for anything about company policy.16- Run independent searches in parallel in a single turn where you can.17- Stop searching once you can answer; do not gather more for its own sake."""1819@dataclass20class Budget:21 max_iterations: int = 1222 max_seconds: float = 180.023 max_input_tokens: int = 400_00024 started: float = field(default_factory=time.monotonic)25 used_input: int = 026 iterations: int = 02728 def exhausted(self) -> str | None:29 if self.iterations >= self.max_iterations:30 return f"iteration limit ({self.max_iterations})"31 if time.monotonic() - self.started > self.max_seconds:32 return f"time limit ({self.max_seconds:.0f}s)"33 if self.used_input > self.max_input_tokens:34 return f"token budget ({self.max_input_tokens:,})"35 return None3637async def research(question: str, manager: MCPManager, state: "RunState") -> str:38 messages = [{"role": "user", "content": question}]39 budget = Budget()4041 while True:42 reason = budget.exhausted()43 if reason:44 messages.append({"role": "user", "content":45 f"Budget reached ({reason}). Write the best answer you can from "46 f"what you already have, and state clearly what remains unverified."})4748 resp = await client.messages.create(49 model=MODEL,50 max_tokens=4096,51 system=SYSTEM,52 tools=manager.tool_defs,53 messages=messages,54 )55 budget.used_input += resp.usage.input_tokens56 budget.iterations += 157 messages.append({"role": "assistant", "content": resp.content})5859 if resp.stop_reason != "tool_use" or reason:60 return "".join(b.text for b in resp.content if b.type == "text")6162 calls = [b for b in resp.content if b.type == "tool_use"]63 results = await asyncio.gather(*[64 execute_one(manager, state, c) for c in calls])65 messages.append({"role": "user", "content": results})6667async def execute_one(manager: MCPManager, state: "RunState", block) -> dict:68 key = hashlib.sha256(69 f"{block.name}:{json.dumps(block.input, sort_keys=True)}".encode()).hexdigest()7071 if key in state.seen_calls:72 text = ("You already ran this exact call and received the result above. "73 "Change the query, try a different source, or answer with what you have.")74 else:75 state.seen_calls.add(key)76 started = time.monotonic()77 text = await manager.call(block.name, block.input)78 state.record(block.name, block.input, text, time.monotonic() - started)7980 return {"type": "tool_result", "tool_use_id": block.id,81 "content": text[:20_000], "is_error": text.startswith(("No such tool",))}The duplicate guard deserves attention, because it is the specific fix for the 19-identical-searches failure. Hashing the tool name plus its canonicalised arguments catches exact repeats, and instead of silently returning the cached result it tells the model that it is repeating itself. Models loop when they cannot tell that an approach has already failed; naming the loop breaks it.
Parallel execution matters too. Three independent searches issued in one turn run concurrently, so the wall-clock cost is the slowest of the three rather than their sum — 1.4 seconds instead of 4.1 — and, more importantly, they cost one model call instead of three.
State the loop must keep
1@dataclass2class Finding:3 tool: str4 query: str5 source_id: str6 excerpt: str7 seconds: float89@dataclass10class RunState:11 question: str12 seen_calls: set[str] = field(default_factory=set)13 findings: list[Finding] = field(default_factory=list)14 errors: list[tuple[str, str]] = field(default_factory=list)1516 def record(self, tool: str, args: dict, text: str, seconds: float) -> None:17 if text.startswith("No such tool") or any(18 marker in text[:80] for marker in ("did not respond", "failed:")):19 self.errors.append((tool, text[:200]))20 return21 for source_id, excerpt in extract_citations(text):22 self.findings.append(Finding(tool, json.dumps(args)[:200],23 source_id, excerpt[:400], seconds))2425 def sources(self) -> list[str]:26 seen, out = set(), []27 for f in self.findings: # preserve discovery order28 if f.source_id not in seen:29 seen.add(f.source_id)30 out.append(f.source_id)31 return outKeeping findings outside the message list is not redundancy. The message list is what the model sees and it will be truncated, summarised or dropped when it grows; the findings ledger is what you use to verify citations, render a sources section, and detect the failure where the agent produced a confident answer from zero successful tool calls. That check is one line and it catches the worst class of bug in the whole system:
1if not state.findings and state.errors:2 return ("I could not retrieve any sources for this question. "3 f"{len(state.errors)} tool call(s) failed: "4 + "; ".join(f"{t}: {msg}" for t, msg in state.errors[:3]))A mock server for development
Developing against live servers is slow, costs quota, and cannot reproduce failures on demand. A mock server that speaks real MCP over stdio fixes all three, and it takes about thirty lines.
1from mcp.server.mcpserver import MCPServer2import os, asyncio, random34mcp = MCPServer("mock-docs")5FAIL_MODE = os.environ.get("MOCK_FAIL", "none") # none|timeout|error|empty|flaky67CORPUS = {8 "POL-114": "Refund policy v4, effective 2026-07-01. Refunds are available "9 "within 30 days of delivery, reduced from 60 days in v3.",10 "POL-098": "Refund policy v3, effective 2025-01-01. Refunds within 60 days.",11 "ENG-421": "Runbook: refund reconciliation job runs nightly at 02:00 UTC.",12}1314@mcp.tool()15async def semantic_search(query: str, k: int = 3) -> str:16 """Search internal documents. Returns document IDs and excerpts."""17 if FAIL_MODE == "timeout":18 await asyncio.sleep(120)19 if FAIL_MODE == "error":20 raise RuntimeError("simulated backend failure")21 if FAIL_MODE == "empty":22 return "No documents matched that query."23 if FAIL_MODE == "flaky" and random.random() < 0.4:24 raise RuntimeError("simulated transient failure")2526 hits = [(doc_id, text) for doc_id, text in CORPUS.items()27 if any(w in text.lower() for w in query.lower().split())][:k]28 if not hits:29 return "No documents matched that query."30 return "\n\n".join(f"[{doc_id}] {text}" for doc_id, text in hits)3132if __name__ == "__main__":33 mcp.run()Because it is a real MCP server, swapping it in is a one-line config change and the whole path — version detection, tool listing, JSON-RPC framing, argument validation — is exercised exactly as in production. That is what distinguishes a mock server from a mocked client object: patching manager.call in a test proves nothing about whether your schemas are valid.
Two traces
The happy path
| Step | What happens | Elapsed | Input tokens |
|---|---|---|---|
| 1 | Model receives the question and 12 tool definitions | 0.0 s | 4,900 |
| 2 | Emits three tool calls in one turn: docs, slack, web | 2.1 s | — |
| 3 | All three run concurrently; slowest is web at 1.4 s | 3.5 s | — |
| 4 | Model reads 3 results (about 3,100 tokens total) | 3.5 s | 8,000 |
| 5 | Sees POL-114 supersedes POL-098; fetches the Slack thread | 5.2 s | — |
| 6 | Thread returns 14 messages, about 900 tokens | 6.0 s | — |
| 7 | Writes the answer with four citations | 9.4 s | 10,300 |
Three model calls, four tool calls, about 23,200 input tokens in total — roughly 0.07 dollars at the same illustrative price, against the 5.70 dollars the unguarded version spent. The saving comes from three things and none of them is a better model: a smaller tool menu, parallel calls collapsing three round trips into one, and a loop that stops when it has enough.
Error recovery
The same question, but the Slack server's token has expired.
iter 1 model -> docs__semantic_search("refund policy change 2026") slack__search_messages("refund policy") docs -> [POL-114] Refund policy v4 ... 30 days, reduced from 60. slack -> slack__search_messages failed: The Slack token is invalid or revoked. Reconnect the workspace.iter 2 model -> docs__semantic_search("refund policy discussion approval") docs -> [POL-114] ... (same document, no discussion thread)iter 3 model -> final answer: "The refund window changed from 60 days to 30 days, effective 2026-07-01 (source: POL-114, superseding POL-098). I could not determine who discussed the change: the Slack integration is disconnected and needs to be reauthorised."That is the behaviour you are engineering for. The agent lost one source, tried a reasonable alternative, and then said what it could not determine and why rather than filling the gap. Two things made it possible: the manager returned the vendor's real reason as readable text instead of raising, and the system prompt said "if the tools do not support a claim, say so". Remove either and the same run produces a confident, invented answer about who approved the policy.
Where these agents go wrong
| Failure | What you see | Root cause | Fix |
|---|---|---|---|
| Tool-name collision | Results from the wrong system | Flat name map across servers | Prefix by alias; raise on duplicates |
| Infinite loop | Same call repeated with tiny variations | No memory of attempts | Hash calls; tell the model it repeated |
| Context overflow mid-run | Failure at iteration 8, never iteration 1 | Unbounded tool results | Truncate results; cap tool count; budget tokens |
| Confident answer, zero sources | Fluent text, no citations | Errors swallowed into empty results | Check the findings ledger before returning |
| One dead server kills the run | Startup crash | All servers treated as required | Mark most as optional; tell the model what is missing |
| Serial when it could be parallel | Four-minute latency, low token use | One call per turn | Prompt for parallel calls; gather them |
| Stale cached tool list | Calls a tool the server removed | Ignoring list_changed | Re-list on the notification |
Every guard in an agent loop exists because a model, given no reason to stop, will not stop. Budgets are not pessimism; they are the control system.
What this means when you build one
The instinct when an agent underperforms is to improve the prompt. In practice the highest-leverage changes are almost always structural, and they are boring.
Curate the tool list ruthlessly. Twelve well-described tools beat forty-seven, on both cost and accuracy, and the curation is a config file rather than a research problem. If two tools would be chosen by the same sentence from a user, merge them or delete one.
Make the loop observable before you make it clever. Log every iteration with the tool name, argument hash, duration, result size and running token count. The 63-call disaster was diagnosable in ninety seconds once those lines existed and essentially undiagnosable before.
Treat every tool result as untrusted text that will be read by a model. Truncate it, and remember that a document fetched from the web can contain instructions aimed at your agent. The tool-result path is the main injection surface in a system like this, which is a strong argument for keeping high-privilege write tools out of any loop that also reads the open web.
Decide what "done" means, explicitly. The loop above stops on three budgets and on the model declining to call a tool. That is a policy, and writing it down as one is the difference between an agent that answers in nine seconds and one that researches until someone kills the process.