Course Content
Model Context Protocol (MCP)
3 sections · 8 lessons
MCP Integration — Building a Server End-to-End
An engineer ships an MCP server for the company's order database. It works. She tests it by hand, the JSON comes back correctly, the process starts cleanly. Then she connects it to the assistant and watches a week of transcripts. The tool is called eleven times in five days, and seven of those calls fail.
The transcripts show why. Her tool is named query, and its description is "Queries the database." The model is holding fourteen tools, three of which sound like they might search for something. When it does pick hers, it passes {"q": "orders from last week"} — plain English into a field the server hands straight to PostgreSQL. And on the two occasions the call succeeded, the server returned all 4,300 matching rows as one JSON blob, roughly 90,000 tokens, which blew the context window and killed the conversation.
Nothing in that story is a protocol bug. Every message was valid JSON-RPC. The failure is that a server is not really a program that answers requests — it is a contract written for a reader who cannot ask clarifying questions. Getting the wire format right is the easy half. This is about both halves.
The wire format underneath everything
MCP is JSON-RPC 2.0. There are exactly three message shapes, and knowing them cold makes every debugging session shorter.
Request
1{2 "jsonrpc": "2.0",3 "id": 42,4 "method": "tools/call",5 "params": {6 "name": "search_orders",7 "arguments": { "customer_email": "amara@example.com", "since": "2026-08-01" }8 }9}The id must be a string or a number, must not be null, and must not match any other request the client still has in flight. It is how the client matches a reply to a call, which matters because responses may arrive out of order — a slow tools/call and a fast tools/list issued together will come back in the wrong sequence, and that is legal. (As explained in the previous lesson, a real request also carries a _meta block with the protocol version and client capabilities, and a real result carries "resultType": "complete"; the examples here leave them out.)
Success response
1{2 "jsonrpc": "2.0",3 "id": 42,4 "result": {5 "content": [6 { "type": "text", "text": "3 orders found for amara@example.com since 2026-08-01." }7 ],8 "structuredContent": {9 "orders": [10 { "id": "ord_8812", "total_cents": 4599, "status": "shipped" },11 { "id": "ord_8907", "total_cents": 12040, "status": "processing" },12 { "id": "ord_9014", "total_cents": 799, "status": "cancelled" }13 ]14 },15 "isError": false16 }17}Error response
1{2 "jsonrpc": "2.0",3 "id": 42,4 "error": {5 "code": -32602,6 "message": "Invalid params",7 "data": { "field": "since", "reason": "expected ISO-8601 date, got 'last week'" }8 }9}A response carries result or error, never both. And there is a fourth shape that is not a message pair at all: a notification, which is a request with no id and therefore no response — used for progress updates, log messages, cancellations on stdio, and list-changed signals.
| Code | Meaning | Typical cause in an MCP server |
|---|---|---|
-32700 | Parse error | A stray print() polluted stdout on the stdio transport |
-32600 | Invalid request | Missing jsonrpc field, or an id of null |
-32601 | Method not found | Client called a method the server never advertised in capabilities |
-32602 | Invalid params | Arguments failed the tool's inputSchema |
-32603 | Internal error | An unhandled exception inside a handler |
Protocol errors are for the developer. Tool failures are for the model, and belong in a normal result with
isError: trueso the agent can read the message and try something else.
Resources: addressable, side-effect-free context
A resource is something the host can read to fill the model's context. It is addressed by URI, it is a GET, and reading it twice must not change anything.
The client discovers what exists with resources/list:
1{2 "jsonrpc": "2.0",3 "id": 3,4 "result": {5 "resources": [6 { "uri": "db://analytics/schema/orders",7 "name": "orders schema",8 "description": "Columns, types and foreign keys for public.orders",9 "mimeType": "application/json" },10 { "uri": "db://analytics/docs/metrics-glossary",11 "name": "Metrics glossary",12 "description": "Definitions of GMV, net revenue, contribution margin",13 "mimeType": "text/markdown" }14 ],15 "nextCursor": "eyJwYWdlIjoyfQ==",16 "ttlMs": 300000,17 "cacheScope": "private"18 }19}Note nextCursor. Listing is paginated: an opaque token the client passes back as params.cursor to get the next page. If it is absent, the list is complete. A server holding 50,000 files must paginate rather than serialise all of them into one response. Note too ttlMs and cacheScope, which 2026-07-28 requires on list and read results: this list may be cached for five minutes, and only by this client, not by a shared proxy.
Reading takes the URI and returns contents:
1{2 "jsonrpc": "2.0",3 "id": 4,4 "method": "resources/read",5 "params": { "uri": "db://analytics/schema/orders" }6}1{2 "jsonrpc": "2.0",3 "id": 4,4 "result": {5 "contents": [6 { "uri": "db://analytics/schema/orders",7 "mimeType": "application/json",8 "text": "{\"columns\":[{\"name\":\"id\",\"type\":\"uuid\"},{\"name\":\"total_cents\",\"type\":\"integer\"}]}" }9 ]10 }11}contents is an array because one URI may legitimately expand to several documents — a directory URI, for instance. Each entry carries either text or blob (base64) but not both.
Templates, for spaces too large to enumerate
A database has thousands of tables. Listing every schema as a static resource is absurd. resources/templates/list returns RFC 6570 URI templates instead:
1{2 "uriTemplate": "db://analytics/schema/{table}",3 "name": "Table schema",4 "description": "Schema for any table in the analytics database",5 "mimeType": "application/json"6}The client fills in {table} and reads. Servers that support it can also implement completion/complete, which lets the host offer autocomplete for that {table} argument as the user types.
Subscriptions
If the server advertised "resources": { "subscribe": true }, a client may ask to be told when a resource changes. In 2026-07-28 it does so by opening a long-lived subscriptions/listen request that names exactly the notifications it wants:
1{2 "jsonrpc": "2.0",3 "id": 9,4 "method": "subscriptions/listen",5 "params": {6 "notifications": {7 "toolsListChanged": true,8 "resourceSubscriptions": ["db://analytics/schema/orders"]9 }10 }11}The server first confirms with notifications/subscriptions/acknowledged, then keeps the response stream open. When the underlying data changes it pushes a notification tagged with the subscription's ID, which is the id of the listen request:
1{2 "jsonrpc": "2.0",3 "method": "notifications/resources/updated",4 "params": {5 "_meta": { "io.modelcontextprotocol/subscriptionId": 9 },6 "uri": "db://analytics/schema/orders"7 }8}Older servers (2025-11-25 and earlier) used a separate resources/subscribe method instead, and sent the notification over the session's stream.
The notification carries no content — only the URI. The client re-reads if it cares. This keeps push messages small and avoids sending 40 KB of JSON to a client that has moved on.
Tools: the model-facing API surface
Tools are where the design work is. A tool definition has a name, a human-readable description, a JSON Schema for its input, optionally a schema for its output, and optionally behavioural annotations.
1{2 "name": "search_orders",3 "title": "Search customer orders",4 "description": "Find orders matching a customer email and/or a date range. Returns at most 50 orders, newest first, each with id, total in cents, status and placed_at. Use get_order_detail for line items. Does not search refunds.",5 "inputSchema": {6 "type": "object",7 "properties": {8 "customer_email": { "type": "string", "format": "email" },9 "since": { "type": "string", "format": "date",10 "description": "ISO-8601 date, e.g. 2026-08-01. Inclusive." },11 "status": { "type": "string",12 "enum": ["processing", "shipped", "delivered", "cancelled"] },13 "limit": { "type": "integer", "minimum": 1, "maximum": 50, "default": 20 }14 },15 "required": ["customer_email"],16 "additionalProperties": false17 },18 "annotations": {19 "readOnlyHint": true,20 "idempotentHint": true,21 "openWorldHint": false22 }23}Compare that with {"name": "query", "description": "Queries the database."} and the earlier failures explain themselves. The rewritten description states the return shape, the cap, the ordering, the neighbouring tool, and one thing it explicitly does not do. The enum makes an invalid status unrepresentable. format: "date" tells the model the shape it must produce, so "last week" never gets sent. additionalProperties: false means a hallucinated argument is rejected at validation rather than silently ignored.
Annotations, and why they are hints
| Annotation | Says | What a host does with it |
|---|---|---|
readOnlyHint | No side effects | May auto-approve without prompting the user |
destructiveHint | May delete or overwrite | Requires explicit confirmation, shows a warning |
idempotentHint | Repeating is harmless | Safe to retry after a timeout |
openWorldHint | Touches external systems | Expect variable latency and network errors |
They are called hints for a reason. They come from the server, and a server is not a trusted authority about its own safety. A host must treat readOnlyHint: true from an unvetted server as a claim, not a guarantee. Where they earn their keep is in reducing consent fatigue for servers the user has already decided to trust.
Calling a tool
1{2 "jsonrpc": "2.0",3 "id": 11,4 "method": "tools/call",5 "params": {6 "name": "search_orders",7 "arguments": { "customer_email": "amara@example.com", "since": "2026-08-01", "limit": 5 }8 }9}The result's content array can hold text, images, audio, or embedded resources — so a chart tool can return a PNG and a document tool can return a resource link the host may fetch later. When a tool declares an outputSchema, it must also return structuredContent conforming to it, and by convention repeat a readable rendering in content for models and hosts that only handle text.
Prompts: parameterised workflows the user triggers
Prompts move prompt engineering to the team that owns the domain. The server publishes a template; the host surfaces it as a slash command.
1{2 "name": "investigate_refund_spike",3 "title": "Investigate a refund spike",4 "description": "Assemble the standard first-pass analysis for a refund anomaly.",5 "arguments": [6 { "name": "region", "description": "Two-letter region code, e.g. DE", "required": true },7 { "name": "window_days", "description": "Lookback window in days", "required": false }8 ]9}prompts/get expands it into actual messages, and those messages may embed resources directly, so the glossary arrives with the instruction rather than being requested separately:
1{2 "jsonrpc": "2.0",3 "id": 21,4 "result": {5 "description": "Refund spike investigation for DE over 14 days",6 "messages": [7 { "role": "user",8 "content": { "type": "text",9 "text": "Refunds in region DE rose over the last 14 days. Compare against the prior 14 days, break the change down by product category and refund reason, and name the two largest contributors." } },10 { "role": "user",11 "content": { "type": "resource",12 "resource": { "uri": "db://analytics/docs/metrics-glossary",13 "mimeType": "text/markdown",14 "text": "Net refund rate = refunded_cents / gross_cents ..." } } }15 ]16 }17}A complete server, end to end
Here is the whole thing in Python, using the official SDK's high-level decorators (version 2 of the mcp package, which speaks 2026-07-28 and falls back to older revisions automatically). It exposes one resource template, one tool and one prompt, and runs over stdio.
1import asyncio, json, re2from datetime import date3from typing import Annotated, Literal4import asyncpg5from pydantic import Field6from mcp.server.mcpserver import MCPServer78mcp = MCPServer("orders-server")9POOL: asyncpg.Pool | None = None1011SAFE_TABLE = re.compile(r"^[a-z_][a-z0-9_]{0,62}$")1213@mcp.resource("db://analytics/schema/{table}")14async def table_schema(table: str) -> str:15 """Column names and types for a table in the analytics database."""16 if not SAFE_TABLE.match(table):17 raise ValueError(f"invalid table name: {table!r}")18 rows = await POOL.fetch(19 """select column_name, data_type, is_nullable20 from information_schema.columns21 where table_schema = 'public' and table_name = $122 order by ordinal_position""",23 table,24 )25 if not rows:26 raise ValueError(f"no such table: {table}")27 return json.dumps({"table": table, "columns": [dict(r) for r in rows]}, indent=2)2829@mcp.tool()30async def search_orders(31 customer_email: Annotated[str, Field(description="The customer's email address.")],32 since: Annotated[str | None, Field(description="ISO-8601 date, e.g. 2026-08-01. Inclusive.")] = None,33 status: Annotated[Literal["processing", "shipped", "delivered", "cancelled"] | None,34 Field(description="Only orders in this status.")] = None,35 limit: Annotated[int, Field(ge=1, le=50, description="Maximum orders to return.")] = 20,36) -> str:37 """Find orders for a customer, newest first, at most 50.3839 Returns id, total_cents, status and placed_at for each order.40 Use get_order_detail for line items. Does not search refunds.41 """42 if since:43 try:44 date.fromisoformat(since)45 except ValueError:46 return f"Invalid 'since' value {since!r}. Expected an ISO-8601 date like 2026-08-01."4748 rows = await POOL.fetch(49 """select id, total_cents, status, placed_at50 from public.orders51 where customer_email = $152 and ($2::date is null or placed_at >= $2::date)53 and ($3::text is null or status = $3)54 order by placed_at desc55 limit $4""",56 customer_email, since, status, limit + 1, # one extra row = "there is more"57 )58 if not rows:59 return f"No orders for {customer_email} matching those filters."60 lines = [f"{r['id']} {r['total_cents']/100:.2f} {r['status']} {r['placed_at']:%Y-%m-%d}"61 for r in rows[:limit]]62 out = f"{len(lines)} order(s):\n" + "\n".join(lines)63 if len(rows) > limit:64 out += f"\n\nShowing the newest {limit}; more matched. Narrow with since or status."65 return out6667@mcp.prompt()68def investigate_refund_spike(region: str, window_days: int = 14) -> str:69 """Assemble the standard first-pass analysis for a refund anomaly."""70 return (f"Refunds in region {region} rose over the last {window_days} days. "71 f"Compare against the prior {window_days} days, break the change down by "72 f"product category and refund reason, and name the two largest contributors.")7374async def main():75 global POOL76 POOL = await asyncpg.create_pool(dsn="postgresql://reader@localhost/analytics",77 min_size=1, max_size=5)78 await mcp.run_stdio_async()7980if __name__ == "__main__":81 asyncio.run(main())Four details in that code carry weight. The table name is checked against a regex before it touches the database, a habit worth keeping because an identifier cannot always be a bind parameter and string-formatting a table name is an injection hole. The Annotated types become the tool's JSON Schema: Literal produces the enum, Field(ge=1, le=50) the minimum and maximum, and each description lands on its property, so the SDK rejects limit=500 before your code runs. The since validation returns a readable message rather than raising, so the model gets a correction it can act on. And the result is a compact fixed-width summary, not a dump of raw rows, with a sentence saying when more rows matched.
If you are reading older code, the same class was called FastMCP and imported from mcp.server.fastmcp in version 1 of the SDK (mcp<2), which spoke only the handshake-based revisions. The decorators are the same; the import and a few advanced APIs changed.
The host is told how to launch it with a small config entry:
1{2 "mcpServers": {3 "orders": {4 "command": "uv",5 "args": ["run", "--directory", "/srv/orders-mcp", "python", "server.py"],6 "env": { "PGPASSWORD": "..." }7 }8 }9}For an interactive test loop, the MCP Inspector speaks the protocol and shows every message. Either launch it directly or through the SDK's CLI (uv run mcp dev server.py, which needs the mcp[cli] extra):
npx @modelcontextprotocol/inspector uv run python server.pyDiscovery and introspection
Everything a client knows about a server, it learned by asking. The usual pattern is: optionally call server/discover, then call tools/list, resources/list and prompts/list, following nextCursor until exhausted, and cache the result for as long as its ttlMs allows. Servers should return tools in a stable order, which also keeps the model provider's prompt cache warm.
Caching creates a staleness problem, which is what the change notifications solve. If the server advertised "tools": { "listChanged": true } and the client opened a subscriptions/listen stream with toolsListChanged: true, the server emits this whenever its catalogue shifts — a plugin loaded, a feature flag flipped:
{ "jsonrpc": "2.0", "method": "notifications/tools/list_changed", "params": { "_meta": { "io.modelcontextprotocol/subscriptionId": 9 } } }The client re-lists. This is what makes dynamic capability possible: an agent can gain a tool without restarting. One thing the list may not do in 2026-07-28 is vary by connection: it may depend on the credentials sent with the request, but not on anything an earlier request did.
One more utility deserves a mention. notifications/progress lets a long-running tool report intermediate status on that request's own response stream, provided the caller supplied a progressToken in params._meta:
1{2 "jsonrpc": "2.0",3 "method": "notifications/progress",4 "params": { "progressToken": "scan-77", "progress": 340, "total": 1200,5 "message": "scanning shard 3 of 8" }6}Designing a contract a model can actually use
The token arithmetic decides more of this than people expect. A well-written tool definition costs roughly 300–400 tokens once serialised into the model's schema format. Forty tools is 14,000 tokens consumed before the user has typed anything — on a 200,000-token window that is 7% of the budget spent on a menu, and measurably worse selection accuracy, because the model is choosing from forty near-synonyms.
| Symptom in transcripts | Cause | Fix |
|---|---|---|
| Tool is never chosen | Description does not name the concepts a user would use | Write the description in the user's vocabulary, list what it returns |
| Wrong tool chosen among similar ones | Overlapping scopes, no stated boundaries | State the boundary explicitly: "Does not search refunds" |
| Arguments arrive malformed | Free-string parameters with no format or enum | Use enum, format, minimum/maximum, and per-field descriptions |
| Context window blows up after one call | Unbounded result size | Cap rows server-side, summarise, return a resource URI for the rest |
| Agent loops retrying the same failing call | Error text says "error" and nothing else | Return the reason plus the corrective action in the error message |
| Model invents a tool that does not exist | Catalogue too large and too granular | Merge fine-grained tools behind one with an operation enum |
Write tool descriptions for a competent new colleague who will read them once, cannot ask you a question, and will be judged on whether they picked the right one.
The unbounded-result failure is worth pricing out. A 4,300-row result at roughly 21 tokens per row is about 90,000 tokens. Capping at 50 rows and adding a one-line summary costs about 1,100 tokens — an eighty-fold reduction — and the agent loses nothing, because a model cannot reason usefully over 4,300 rows anyway. When the full set genuinely matters, return a resource URI pointing at the export and let the host decide whether to fetch it.
What this means when you build one
Treat the tool catalogue as a product surface with a user who happens to be a model. That reframing produces a short set of habits worth adopting from the first commit.
- Start from the transcripts, not the schema. Write down the five sentences a user would actually type, then design the smallest set of tools that answers all five. Tools that no plausible sentence reaches are dead weight in every prompt.
- Bound every output at the server. Every list gets a maximum. Every text field gets a truncation. The server, not the agent, is the last place with the information needed to summarise sensibly.
- Make failures instructive. "Invalid date" is a dead end; "expected ISO-8601 like 2026-08-01, got 'last week'" gets a correct retry on the next turn.
- Validate at the boundary and again in the handler. A schema is a description, not enforcement — a non-conforming client can send anything, so range checks and identifier regexes belong in the code too.
- Version by adding, not changing. Renaming a tool or tightening a required field breaks every host that cached your catalogue. Add a new tool and mark the old description as superseded.
The engineer from the opening rewrote her single query tool as three: search_orders, get_order_detail, and order_stats. Each with enums, a hard row cap, and a description naming what it does not cover. Call volume went from eleven in a week to roughly forty a day, and the failure rate fell to under one in twenty. The protocol never changed. The contract did.