Model Context Protocol (MCP)

MCP Integration — Building a Server End-to-End


An engineer ships an MCP server for the company's order database. It works. She tests it by hand, the JSON comes back correctly, the process starts cleanly. Then she connects it to the assistant and watches a week of transcripts. The tool is called eleven times in five days, and seven of those calls fail.

The transcripts show why. Her tool is named query, and its description is "Queries the database." The model is holding fourteen tools, three of which sound like they might search for something. When it does pick hers, it passes {"q": "orders from last week"} — plain English into a field the server hands straight to PostgreSQL. And on the two occasions the call succeeded, the server returned all 4,300 matching rows as one JSON blob, roughly 90,000 tokens, which blew the context window and killed the conversation.

Nothing in that story is a protocol bug. Every message was valid JSON-RPC. The failure is that a server is not really a program that answers requests — it is a contract written for a reader who cannot ask clarifying questions. Getting the wire format right is the easy half. This is about both halves.

What a server puts on the wireMCP serverResources — read,no side effectsTemplates for large spacesSubscriptions on changeTools — themodel-facing verbsPrompts the user triggersDiscovery: list before call
Eleven calls and seven wrong means the contract is the product — descriptions and annotations are read by the model, not by you.

The wire format underneath everything

MCP is JSON-RPC 2.0. There are exactly three message shapes, and knowing them cold makes every debugging session shorter.

Request

JSON
{  "jsonrpc": "2.0",  "id": 42,  "method": "tools/call",  "params": {    "name": "search_orders",    "arguments": { "customer_email": "amara@example.com", "since": "2026-08-01" }  }}

The id must be a string or a number, must not be null, and must not match any other request the client still has in flight. It is how the client matches a reply to a call, which matters because responses may arrive out of order — a slow tools/call and a fast tools/list issued together will come back in the wrong sequence, and that is legal. (As explained in the previous lesson, a real request also carries a _meta block with the protocol version and client capabilities, and a real result carries "resultType": "complete"; the examples here leave them out.)

Success response

JSON
{  "jsonrpc": "2.0",  "id": 42,  "result": {    "content": [      { "type": "text", "text": "3 orders found for amara@example.com since 2026-08-01." }    ],    "structuredContent": {      "orders": [        { "id": "ord_8812", "total_cents": 4599, "status": "shipped" },        { "id": "ord_8907", "total_cents": 12040, "status": "processing" },        { "id": "ord_9014", "total_cents": 799,  "status": "cancelled" }      ]    },    "isError": false  }}

Error response

JSON
{  "jsonrpc": "2.0",  "id": 42,  "error": {    "code": -32602,    "message": "Invalid params",    "data": { "field": "since", "reason": "expected ISO-8601 date, got 'last week'" }  }}

A response carries result or error, never both. And there is a fourth shape that is not a message pair at all: a notification, which is a request with no id and therefore no response — used for progress updates, log messages, cancellations on stdio, and list-changed signals.

CodeMeaningTypical cause in an MCP server
-32700Parse errorA stray print() polluted stdout on the stdio transport
-32600Invalid requestMissing jsonrpc field, or an id of null
-32601Method not foundClient called a method the server never advertised in capabilities
-32602Invalid paramsArguments failed the tool's inputSchema
-32603Internal errorAn unhandled exception inside a handler

Protocol errors are for the developer. Tool failures are for the model, and belong in a normal result with isError: true so the agent can read the message and try something else.

Resources: addressable, side-effect-free context

A resource is something the host can read to fill the model's context. It is addressed by URI, it is a GET, and reading it twice must not change anything.

The client discovers what exists with resources/list:

JSON
{  "jsonrpc": "2.0",  "id": 3,  "result": {    "resources": [      { "uri": "db://analytics/schema/orders",        "name": "orders schema",        "description": "Columns, types and foreign keys for public.orders",        "mimeType": "application/json" },      { "uri": "db://analytics/docs/metrics-glossary",        "name": "Metrics glossary",        "description": "Definitions of GMV, net revenue, contribution margin",        "mimeType": "text/markdown" }    ],    "nextCursor": "eyJwYWdlIjoyfQ==",    "ttlMs": 300000,    "cacheScope": "private"  }}

Note nextCursor. Listing is paginated: an opaque token the client passes back as params.cursor to get the next page. If it is absent, the list is complete. A server holding 50,000 files must paginate rather than serialise all of them into one response. Note too ttlMs and cacheScope, which 2026-07-28 requires on list and read results: this list may be cached for five minutes, and only by this client, not by a shared proxy.

Reading takes the URI and returns contents:

JSON
{  "jsonrpc": "2.0",  "id": 4,  "method": "resources/read",  "params": { "uri": "db://analytics/schema/orders" }}
JSON
{  "jsonrpc": "2.0",  "id": 4,  "result": {    "contents": [      { "uri": "db://analytics/schema/orders",        "mimeType": "application/json",        "text": "{\"columns\":[{\"name\":\"id\",\"type\":\"uuid\"},{\"name\":\"total_cents\",\"type\":\"integer\"}]}" }    ]  }}

contents is an array because one URI may legitimately expand to several documents — a directory URI, for instance. Each entry carries either text or blob (base64) but not both.

Templates, for spaces too large to enumerate

A database has thousands of tables. Listing every schema as a static resource is absurd. resources/templates/list returns RFC 6570 URI templates instead:

JSON
{  "uriTemplate": "db://analytics/schema/{table}",  "name": "Table schema",  "description": "Schema for any table in the analytics database",  "mimeType": "application/json"}

The client fills in {table} and reads. Servers that support it can also implement completion/complete, which lets the host offer autocomplete for that {table} argument as the user types.

Subscriptions

If the server advertised "resources": { "subscribe": true }, a client may ask to be told when a resource changes. In 2026-07-28 it does so by opening a long-lived subscriptions/listen request that names exactly the notifications it wants:

JSON
{  "jsonrpc": "2.0",  "id": 9,  "method": "subscriptions/listen",  "params": {    "notifications": {      "toolsListChanged": true,      "resourceSubscriptions": ["db://analytics/schema/orders"]    }  }}

The server first confirms with notifications/subscriptions/acknowledged, then keeps the response stream open. When the underlying data changes it pushes a notification tagged with the subscription's ID, which is the id of the listen request:

JSON
{  "jsonrpc": "2.0",  "method": "notifications/resources/updated",  "params": {    "_meta": { "io.modelcontextprotocol/subscriptionId": 9 },    "uri": "db://analytics/schema/orders"  }}

Older servers (2025-11-25 and earlier) used a separate resources/subscribe method instead, and sent the notification over the session's stream.

The notification carries no content — only the URI. The client re-reads if it cares. This keeps push messages small and avoids sending 40 KB of JSON to a client that has moved on.

Tools: the model-facing API surface

Tools are where the design work is. A tool definition has a name, a human-readable description, a JSON Schema for its input, optionally a schema for its output, and optionally behavioural annotations.

JSON
{  "name": "search_orders",  "title": "Search customer orders",  "description": "Find orders matching a customer email and/or a date range. Returns at most 50 orders, newest first, each with id, total in cents, status and placed_at. Use get_order_detail for line items. Does not search refunds.",  "inputSchema": {    "type": "object",    "properties": {      "customer_email": { "type": "string", "format": "email" },      "since": { "type": "string", "format": "date",                 "description": "ISO-8601 date, e.g. 2026-08-01. Inclusive." },      "status": { "type": "string",                  "enum": ["processing", "shipped", "delivered", "cancelled"] },      "limit": { "type": "integer", "minimum": 1, "maximum": 50, "default": 20 }    },    "required": ["customer_email"],    "additionalProperties": false  },  "annotations": {    "readOnlyHint": true,    "idempotentHint": true,    "openWorldHint": false  }}

Compare that with {"name": "query", "description": "Queries the database."} and the earlier failures explain themselves. The rewritten description states the return shape, the cap, the ordering, the neighbouring tool, and one thing it explicitly does not do. The enum makes an invalid status unrepresentable. format: "date" tells the model the shape it must produce, so "last week" never gets sent. additionalProperties: false means a hallucinated argument is rejected at validation rather than silently ignored.

Annotations, and why they are hints

AnnotationSaysWhat a host does with it
readOnlyHintNo side effectsMay auto-approve without prompting the user
destructiveHintMay delete or overwriteRequires explicit confirmation, shows a warning
idempotentHintRepeating is harmlessSafe to retry after a timeout
openWorldHintTouches external systemsExpect variable latency and network errors

They are called hints for a reason. They come from the server, and a server is not a trusted authority about its own safety. A host must treat readOnlyHint: true from an unvetted server as a claim, not a guarantee. Where they earn their keep is in reducing consent fatigue for servers the user has already decided to trust.

Calling a tool

JSON
{  "jsonrpc": "2.0",  "id": 11,  "method": "tools/call",  "params": {    "name": "search_orders",    "arguments": { "customer_email": "amara@example.com", "since": "2026-08-01", "limit": 5 }  }}

The result's content array can hold text, images, audio, or embedded resources — so a chart tool can return a PNG and a document tool can return a resource link the host may fetch later. When a tool declares an outputSchema, it must also return structuredContent conforming to it, and by convention repeat a readable rendering in content for models and hosts that only handle text.

Prompts: parameterised workflows the user triggers

Prompts move prompt engineering to the team that owns the domain. The server publishes a template; the host surfaces it as a slash command.

JSON
{  "name": "investigate_refund_spike",  "title": "Investigate a refund spike",  "description": "Assemble the standard first-pass analysis for a refund anomaly.",  "arguments": [    { "name": "region", "description": "Two-letter region code, e.g. DE", "required": true },    { "name": "window_days", "description": "Lookback window in days", "required": false }  ]}

prompts/get expands it into actual messages, and those messages may embed resources directly, so the glossary arrives with the instruction rather than being requested separately:

JSON
{  "jsonrpc": "2.0",  "id": 21,  "result": {    "description": "Refund spike investigation for DE over 14 days",    "messages": [      { "role": "user",        "content": { "type": "text",          "text": "Refunds in region DE rose over the last 14 days. Compare against the prior 14 days, break the change down by product category and refund reason, and name the two largest contributors." } },      { "role": "user",        "content": { "type": "resource",          "resource": { "uri": "db://analytics/docs/metrics-glossary",                        "mimeType": "text/markdown",                        "text": "Net refund rate = refunded_cents / gross_cents ..." } } }    ]  }}

A complete server, end to end

Here is the whole thing in Python, using the official SDK's high-level decorators (version 2 of the mcp package, which speaks 2026-07-28 and falls back to older revisions automatically). It exposes one resource template, one tool and one prompt, and runs over stdio.

Python
import asyncio, json, refrom datetime import datefrom typing import Annotated, Literalimport asyncpgfrom pydantic import Fieldfrom mcp.server.mcpserver import MCPServermcp = MCPServer("orders-server")POOL: asyncpg.Pool | None = NoneSAFE_TABLE = re.compile(r"^[a-z_][a-z0-9_]{0,62}$")@mcp.resource("db://analytics/schema/{table}")async def table_schema(table: str) -> str:    """Column names and types for a table in the analytics database."""    if not SAFE_TABLE.match(table):        raise ValueError(f"invalid table name: {table!r}")    rows = await POOL.fetch(        """select column_name, data_type, is_nullable             from information_schema.columns            where table_schema = 'public' and table_name = $1            order by ordinal_position""",        table,    )    if not rows:        raise ValueError(f"no such table: {table}")    return json.dumps({"table": table, "columns": [dict(r) for r in rows]}, indent=2)@mcp.tool()async def search_orders(    customer_email: Annotated[str, Field(description="The customer's email address.")],    since: Annotated[str | None, Field(description="ISO-8601 date, e.g. 2026-08-01. Inclusive.")] = None,    status: Annotated[Literal["processing", "shipped", "delivered", "cancelled"] | None,                      Field(description="Only orders in this status.")] = None,    limit: Annotated[int, Field(ge=1, le=50, description="Maximum orders to return.")] = 20,) -> str:    """Find orders for a customer, newest first, at most 50.    Returns id, total_cents, status and placed_at for each order.    Use get_order_detail for line items. Does not search refunds.    """    if since:        try:            date.fromisoformat(since)        except ValueError:            return f"Invalid 'since' value {since!r}. Expected an ISO-8601 date like 2026-08-01."    rows = await POOL.fetch(        """select id, total_cents, status, placed_at             from public.orders            where customer_email = $1              and ($2::date is null or placed_at >= $2::date)              and ($3::text is null or status = $3)            order by placed_at desc            limit $4""",        customer_email, since, status, limit + 1,    # one extra row = "there is more"    )    if not rows:        return f"No orders for {customer_email} matching those filters."    lines = [f"{r['id']}  {r['total_cents']/100:.2f}  {r['status']}  {r['placed_at']:%Y-%m-%d}"             for r in rows[:limit]]    out = f"{len(lines)} order(s):\n" + "\n".join(lines)    if len(rows) > limit:        out += f"\n\nShowing the newest {limit}; more matched. Narrow with since or status."    return out@mcp.prompt()def investigate_refund_spike(region: str, window_days: int = 14) -> str:    """Assemble the standard first-pass analysis for a refund anomaly."""    return (f"Refunds in region {region} rose over the last {window_days} days. "            f"Compare against the prior {window_days} days, break the change down by "            f"product category and refund reason, and name the two largest contributors.")async def main():    global POOL    POOL = await asyncpg.create_pool(dsn="postgresql://reader@localhost/analytics",                                     min_size=1, max_size=5)    await mcp.run_stdio_async()if __name__ == "__main__":    asyncio.run(main())

Four details in that code carry weight. The table name is checked against a regex before it touches the database, a habit worth keeping because an identifier cannot always be a bind parameter and string-formatting a table name is an injection hole. The Annotated types become the tool's JSON Schema: Literal produces the enum, Field(ge=1, le=50) the minimum and maximum, and each description lands on its property, so the SDK rejects limit=500 before your code runs. The since validation returns a readable message rather than raising, so the model gets a correction it can act on. And the result is a compact fixed-width summary, not a dump of raw rows, with a sentence saying when more rows matched.

If you are reading older code, the same class was called FastMCP and imported from mcp.server.fastmcp in version 1 of the SDK (mcp<2), which spoke only the handshake-based revisions. The decorators are the same; the import and a few advanced APIs changed.

The host is told how to launch it with a small config entry:

JSON
{  "mcpServers": {    "orders": {      "command": "uv",      "args": ["run", "--directory", "/srv/orders-mcp", "python", "server.py"],      "env": { "PGPASSWORD": "..." }    }  }}

For an interactive test loop, the MCP Inspector speaks the protocol and shows every message. Either launch it directly or through the SDK's CLI (uv run mcp dev server.py, which needs the mcp[cli] extra):

Bash
npx @modelcontextprotocol/inspector uv run python server.py

Discovery and introspection

Everything a client knows about a server, it learned by asking. The usual pattern is: optionally call server/discover, then call tools/list, resources/list and prompts/list, following nextCursor until exhausted, and cache the result for as long as its ttlMs allows. Servers should return tools in a stable order, which also keeps the model provider's prompt cache warm.

Caching creates a staleness problem, which is what the change notifications solve. If the server advertised "tools": { "listChanged": true } and the client opened a subscriptions/listen stream with toolsListChanged: true, the server emits this whenever its catalogue shifts — a plugin loaded, a feature flag flipped:

JSON
{ "jsonrpc": "2.0", "method": "notifications/tools/list_changed",  "params": { "_meta": { "io.modelcontextprotocol/subscriptionId": 9 } } }

The client re-lists. This is what makes dynamic capability possible: an agent can gain a tool without restarting. One thing the list may not do in 2026-07-28 is vary by connection: it may depend on the credentials sent with the request, but not on anything an earlier request did.

One more utility deserves a mention. notifications/progress lets a long-running tool report intermediate status on that request's own response stream, provided the caller supplied a progressToken in params._meta:

JSON
{  "jsonrpc": "2.0",  "method": "notifications/progress",  "params": { "progressToken": "scan-77", "progress": 340, "total": 1200,              "message": "scanning shard 3 of 8" }}

Designing a contract a model can actually use

The token arithmetic decides more of this than people expect. A well-written tool definition costs roughly 300–400 tokens once serialised into the model's schema format. Forty tools is 14,000 tokens consumed before the user has typed anything — on a 200,000-token window that is 7% of the budget spent on a menu, and measurably worse selection accuracy, because the model is choosing from forty near-synonyms.

Symptom in transcriptsCauseFix
Tool is never chosenDescription does not name the concepts a user would useWrite the description in the user's vocabulary, list what it returns
Wrong tool chosen among similar onesOverlapping scopes, no stated boundariesState the boundary explicitly: "Does not search refunds"
Arguments arrive malformedFree-string parameters with no format or enumUse enum, format, minimum/maximum, and per-field descriptions
Context window blows up after one callUnbounded result sizeCap rows server-side, summarise, return a resource URI for the rest
Agent loops retrying the same failing callError text says "error" and nothing elseReturn the reason plus the corrective action in the error message
Model invents a tool that does not existCatalogue too large and too granularMerge fine-grained tools behind one with an operation enum

Write tool descriptions for a competent new colleague who will read them once, cannot ask you a question, and will be judged on whether they picked the right one.

The unbounded-result failure is worth pricing out. A 4,300-row result at roughly 21 tokens per row is about 90,000 tokens. Capping at 50 rows and adding a one-line summary costs about 1,100 tokens — an eighty-fold reduction — and the agent loses nothing, because a model cannot reason usefully over 4,300 rows anyway. When the full set genuinely matters, return a resource URI pointing at the export and let the host decide whether to fetch it.

What this means when you build one

Treat the tool catalogue as a product surface with a user who happens to be a model. That reframing produces a short set of habits worth adopting from the first commit.

  • Start from the transcripts, not the schema. Write down the five sentences a user would actually type, then design the smallest set of tools that answers all five. Tools that no plausible sentence reaches are dead weight in every prompt.
  • Bound every output at the server. Every list gets a maximum. Every text field gets a truncation. The server, not the agent, is the last place with the information needed to summarise sensibly.
  • Make failures instructive. "Invalid date" is a dead end; "expected ISO-8601 like 2026-08-01, got 'last week'" gets a correct retry on the next turn.
  • Validate at the boundary and again in the handler. A schema is a description, not enforcement — a non-conforming client can send anything, so range checks and identifier regexes belong in the code too.
  • Version by adding, not changing. Renaming a tool or tightening a required field breaks every host that cached your catalogue. Add a new tool and mark the old description as superseded.

The engineer from the opening rewrote her single query tool as three: search_orders, get_order_detail, and order_stats. Each with enums, a hard row cap, and a description naming what it does not cover. Call volume went from eleven in a week to roughly forty a day, and the failure rate fell to under one in twenty. The protocol never changed. The contract did.