Course Content
Model Context Protocol (MCP)
3 sections · 8 lessons
Integrating External APIs (Google, Slack, and Beyond)
At 14:50 the Slack MCP server is fine. At 15:00, when the second standup ends and four analysts start asking the assistant questions at once, every slack_search call starts coming back empty. No exception, no error in the agent's transcript — the tool returns "No messages found" and the model confidently tells three people that the incident was never discussed.
Two things went wrong, and both are typical. First, the arithmetic: search.messages sits in a Slack rate-limit tier allowing roughly 20 requests per minute per workspace. Each analyst question fans out to eight searches. Four analysts asking one question a minute is 32 requests a minute against a ceiling of 20, so twelve requests are rejected every minute with HTTP 429 and a Retry-After: 30 header.
Second, and worse: the server retried those twelve after exactly 30 seconds — all twelve at the same instant, because they all failed at the same instant and all waited the same amount. The retry burst collided with the next minute's fresh traffic and the failure rate went up, not down. Meanwhile the handler treated a non-200 response as "no results" and returned a cheerful empty list, so nothing anywhere in the system said the word "failing".
An MCP server that wraps a third-party API is not a thin proxy. It is the component that turns an unreliable, rate-limited, paginated, idiosyncratic HTTP API into something a language model can use without lying to its user.
What actually belongs in the server
| Layer | Job | What breaks without it |
|---|---|---|
| Credential management | Hold keys and refresh tokens; never surface them as tool arguments | The model is asked to supply a secret, and eventually invents one |
| Client-side rate limiting | Stay under the quota before sending | You discover the limit by being throttled |
| Retry with backoff and jitter | Recover from transient failures without synchronised bursts | Retry storms turn a blip into an outage |
| Error mapping | Turn HTTP and vendor error codes into sentences the model can act on | "Error 400" produces an identical retry, forever |
| Pagination | Walk cursors with a hard page cap | One call fetches 40,000 messages into the context window |
| Normalisation | Strip vendor payloads to the fields that matter | 90% of the tokens are metadata the model ignores |
The model should never see an HTTP status code, a cursor token, or a vendor-specific error string. Those are the server's problem, and hiding them is most of its value.
Authentication patterns
API keys: for APIs that identify an application
Google Custom Search, most weather APIs, most search APIs. One key, loaded from the environment, applied to every request. The rule that matters: the key never appears in the tool's input schema. If api_key is a parameter, the model will eventually be asked to fill it in, and a model asked for a credential will hallucinate a plausible-looking string.
1import os, httpx23GOOGLE_KEY = os.environ["GOOGLE_API_KEY"] # fail fast at startup, not mid-call4GOOGLE_CX = os.environ["GOOGLE_CSE_ID"]56client = httpx.AsyncClient(7 timeout=httpx.Timeout(connect=3.0, read=15.0, write=5.0, pool=5.0),8 limits=httpx.Limits(max_connections=20, max_keepalive_connections=10),9)Note the split timeouts. A single timeout=30 means a dead host ties up a worker for thirty seconds; a three-second connect timeout fails fast on unreachable hosts while still allowing a slow search fifteen seconds to think.
OAuth 2.0: for APIs that act on behalf of a person
Slack, Google Drive, GitHub, Notion. The server holds a long-lived refresh token and exchanges it for short-lived access tokens. The part people get wrong is refreshing reactively — waiting for a 401 and then refreshing — which means every session pays at least one failed request, and concurrent calls all try to refresh at once.
1import asyncio, time23class TokenStore:4 """One access token per user, refreshed proactively under a lock."""56 def __init__(self):7 self._tokens: dict[str, tuple[str, float]] = {} # user -> (token, expires_at)8 self._locks: dict[str, asyncio.Lock] = {}910 async def get(self, user_id: str) -> str:11 token, expires_at = self._tokens.get(user_id, (None, 0.0))12 # Refresh 60 s early so an in-flight request never expires mid-call.13 if token and time.time() < expires_at - 60:14 return token1516 lock = self._locks.setdefault(user_id, asyncio.Lock())17 async with lock:18 # Re-check: another coroutine may have refreshed while we waited.19 token, expires_at = self._tokens.get(user_id, (None, 0.0))20 if token and time.time() < expires_at - 60:21 return token22 token, ttl = await self._refresh(user_id)23 self._tokens[user_id] = (token, time.time() + ttl)24 return token2526 async def _refresh(self, user_id: str) -> tuple[str, float]:27 r = await client.post("https://slack.com/api/oauth.v2.access", data={28 "grant_type": "refresh_token",29 "refresh_token": await load_refresh_token(user_id), # from encrypted store30 "client_id": os.environ["SLACK_CLIENT_ID"],31 "client_secret": os.environ["SLACK_CLIENT_SECRET"],32 })33 body = r.json()34 if not body.get("ok"):35 raise AuthError(f"token refresh failed: {body.get('error')}")36 await save_refresh_token(user_id, body["refresh_token"]) # rotation37 return body["access_token"], float(body.get("expires_in", 43200))Three details earn their place. The 60-second early refresh removes the race where a token expires between the check and the request arriving at the server. The double-checked lock stops ten concurrent tool calls from firing ten refreshes — which with rotating refresh tokens invalidates nine of them and locks the user out. And save_refresh_token runs on every refresh because many providers rotate the refresh token too; ignore the new one and the integration silently dies the next time the old one is used.
For a multi-user server, the identity comes from the access token that arrives with each MCP request, never from a tool argument. A tool called read_slack_as(user_id) is an impersonation API with extra steps.
Rate limits and retries are one problem
People implement these separately and then discover they fight each other. A retry policy without a rate limiter reacts to throttling by generating more traffic. A rate limiter without retries drops legitimate work on the first transient blip.
The limiter: stay under the ceiling by construction
1import asyncio, time23class AsyncRateLimiter:4 """Token bucket. Callers await capacity instead of being rejected."""56 def __init__(self, rate_per_sec: float, burst: float):7 self.rate, self.burst = rate_per_sec, burst8 self.tokens, self.updated = burst, time.monotonic()9 self._lock = asyncio.Lock()1011 async def acquire(self, cost: float = 1.0) -> None:12 async with self._lock:13 while True:14 now = time.monotonic()15 self.tokens = min(self.burst,16 self.tokens + (now - self.updated) * self.rate)17 self.updated = now18 if self.tokens >= cost:19 self.tokens -= cost20 return21 await asyncio.sleep((cost - self.tokens) / self.rate)2223# Slack tier-2 methods: ~20 per minute. Leave headroom.24slack_search_limiter = AsyncRateLimiter(rate_per_sec=18 / 60, burst=5)Set the rate slightly below the documented ceiling — 18 per minute against a limit of 20 — because your clock and the vendor's window boundaries do not line up, and running at exactly the limit guarantees occasional 429s.
The retry: back off, and add jitter
Exponential backoff doubles the wait after each failure: 1 s, 2 s, 4 s, 8 s, 16 s. Five attempts span about 31 seconds of waiting. That is the right shape, and on its own it is what caused the outage in the opening — twelve simultaneous failures produce twelve simultaneous retries.
Full jitter fixes it by waiting a uniformly random time between zero and the exponential ceiling. The expected wait is halved, and twelve retries spread themselves across the whole window instead of stacking on one instant.
| Attempt | Fixed 30 s | Exponential, no jitter | Full jitter: random(0, 2^n) |
|---|---|---|---|
| 1 | 30.0 s | 1.0 s | uniform in [0, 1] s |
| 2 | 30.0 s | 2.0 s | uniform in [0, 2] s |
| 3 | 30.0 s | 4.0 s | uniform in [0, 4] s |
| 4 | 30.0 s | 8.0 s | uniform in [0, 8] s |
| 5 | 30.0 s | 16.0 s | uniform in [0, 16] s |
| Collision behaviour | All retries land together | All retries land together | Spread across the window |
1import random2from tenacity import (retry, stop_after_attempt, wait_random_exponential,3 retry_if_exception_type)45class Retryable(Exception):6 """Transient. Worth another attempt."""7 def __init__(self, msg, retry_after: float | None = None):8 super().__init__(msg)9 self.retry_after = retry_after1011class Permanent(Exception):12 """Retrying will produce the identical failure."""1314@retry(retry=retry_if_exception_type(Retryable),15 wait=wait_random_exponential(multiplier=1, max=16), # full jitter16 stop=stop_after_attempt(5),17 reraise=True)18async def call_api(method: str, url: str, **kw) -> dict:19 r = await client.request(method, url, **kw)2021 if r.status_code == 429:22 # Always prefer the server's own advice, then add jitter on top.23 wait = float(r.headers.get("Retry-After", 5))24 await asyncio.sleep(wait + random.uniform(0, 1.0))25 raise Retryable(f"rate limited, server asked for {wait}s", retry_after=wait)2627 if 500 <= r.status_code < 600:28 raise Retryable(f"upstream {r.status_code}")2930 if r.status_code in (401, 403):31 raise Permanent(f"not authorised ({r.status_code}) - check scopes: {r.text[:200]}")3233 if 400 <= r.status_code < 500:34 raise Permanent(f"request rejected ({r.status_code}): {r.text[:200]}")3536 return r.json()Which failures are worth retrying
| Response | Retry? | Why | What the model should be told |
|---|---|---|---|
| 429 Too Many Requests | Yes, after Retry-After | Purely temporal | Only after retries are exhausted: "busy, try again shortly" |
| 500 / 502 / 503 / 504 | Yes, with backoff | Upstream instability | "The service is unavailable right now" |
| Connection reset / timeout | Yes, if the call is idempotent | Network, not logic | Nothing, if a retry succeeds |
| 401 Unauthorized | Once, after refreshing the token | Usually an expired token | "Reconnect your account" |
| 403 Forbidden | No | A missing scope will still be missing | "This needs the files:read scope" |
| 404 Not Found | No | The object does not exist | "No channel named #incidnets — did you mean #incidents?" |
| 400 Bad Request | No | The arguments are wrong | The vendor's validation message, verbatim |
The distinction is the whole point of the two exception classes. Retrying a 403 five times with backoff wastes 31 seconds to arrive at the same 403, and the agent's user watches a spinner for half a minute to learn nothing.
A rate limiter without retries drops legitimate work; retries without a rate limiter answer throttling by sending more traffic. Ship them together or ship neither.
A Google search server
Google's Custom Search JSON API returns at most 10 results per request and at most 100 in total, and the free tier is 100 queries per day. Those constraints belong in the tool description, because a model that does not know about them will paginate happily into a wall.
1from mcp.server.mcpserver import MCPServer23mcp = MCPServer("web-search")4google_limiter = AsyncRateLimiter(rate_per_sec=1.0, burst=3)56@mcp.tool()7async def web_search(query: str, num_results: int = 5,8 site: str | None = None) -> str:9 """Search the public web and return titles, URLs and snippets.1011 Returns at most 10 results per call, newest ranking first.12 Use `site` to restrict to one domain, e.g. site="docs.python.org".13 Snippets are ~150 characters; fetch the URL for full text.14 """15 num_results = max(1, min(num_results, 10))16 q = f"site:{site} {query}" if site else query1718 await google_limiter.acquire()19 try:20 data = await call_api("GET", "https://www.googleapis.com/customsearch/v1",21 params={"key": GOOGLE_KEY, "cx": GOOGLE_CX,22 "q": q, "num": num_results})23 except Permanent as e:24 if "quota" in str(e).lower(): # Google reports quota as a 40325 return ("The daily search quota is exhausted. Answer from what you "26 "already have, or ask the user to try tomorrow.")27 return f"Search failed and retrying will not help: {e}"28 except Retryable as e:29 return f"Search is temporarily unavailable ({e}). Try again in a minute."3031 items = data.get("items", [])32 if not items:33 total = data.get("searchInformation", {}).get("totalResults", "0")34 return (f"No results for {q!r} (Google reported {total} total). "35 f"Try broader terms or drop the site restriction.")3637 return "\n\n".join(38 f"{i}. {it['title']}\n {it['link']}\n {it.get('snippet','').strip()}"39 for i, it in enumerate(items, 1))The quota message is the interesting line. "Quota exceeded" tells the model nothing useful; "answer from what you already have, or ask the user to try tomorrow" tells it what to do instead of retrying.
A Slack server, and Slack's favourite trap
Slack returns HTTP 200 with {"ok": false, "error": "..."} for most application-level failures. A handler that checks only r.status_code == 200 treats "channel_not_found" and "missing_scope" as success and returns an empty result — which is precisely the bug that let the agent tell three people the incident was never discussed.
1SLACK_ERRORS = {2 "channel_not_found": "That channel does not exist or the app is not a member of it.",3 "not_in_channel": "The app must be invited to that channel first (/invite @bot).",4 "missing_scope": "The Slack app is missing a required OAuth scope.",5 "invalid_auth": "The Slack token is invalid or revoked. Reconnect the workspace.",6 "ratelimited": "Slack is throttling this workspace. Wait a moment.",7}89async def slack_call(method: str, user_id: str, **params) -> dict:10 await slack_search_limiter.acquire()11 token = await tokens.get(user_id)12 body = await call_api("POST", f"https://slack.com/api/{method}",13 headers={"Authorization": f"Bearer {token}"},14 data=params)15 if not body.get("ok"):16 err = body.get("error", "unknown_error")17 raise Permanent(SLACK_ERRORS.get(err, f"Slack rejected the call: {err}"))18 return body1920from mcp.server.auth.middleware.auth_context import get_access_token2122@mcp.tool()23async def slack_search(query: str, count: int = 10) -> str:24 """Search messages across channels the app can see.2526 Supports Slack search syntax: in:#channel, from:@user, before:2026-08-01.27 Returns at most 20 messages, most relevant first, with permalinks.28 """29 count = max(1, min(count, 20))30 access = get_access_token() # the caller's validated OAuth token, per request31 if access is None or access.subject is None:32 return "This tool needs a signed-in user. Ask the user to connect their account."33 try:34 body = await slack_call("search.messages", access.subject,35 query=query, count=count, sort="score")36 except Permanent as e:37 return f"Slack search failed: {e}"3839 matches = body["messages"]["matches"]40 total = body["messages"]["total"]41 if not matches:42 return (f"No messages matched {query!r}. Slack search only covers channels "43 f"the app has joined - try naming a channel with in:#name.")4445 lines = [f"{total} match(es); showing {len(matches)}:"]46 for m in matches:47 text = " ".join(m.get("text", "").split())[:400]48 lines.append(f"[#{m['channel']['name']} {m['username']} "49 f"{m['ts'][:10]}]\n{text}\n{m['permalink']}")50 return "\n\n".join(lines)The user's identity comes from get_access_token(), which the SDK fills in from the bearer token it validated for this request; the model never sees or supplies it. The empty-result message does real work: it explains why a search might legitimately find nothing and suggests a narrower query. That single sentence converts a dead end into a next step.
Pagination
Three schemes dominate, and each fails differently.
| Scheme | Looks like | Used by | Failure mode |
|---|---|---|---|
| Cursor | next_cursor in the response | Slack, Stripe | Cursors expire; resuming an old one 400s |
| Offset / page | start=11, page=3 | Google CSE, older REST APIs | Rows inserted mid-walk cause duplicates and skips |
| Link header | Link: <...>; rel="next" | GitHub | Easy to parse wrongly; the URL must be used verbatim |
1async def paginate_slack(method: str, user_id: str, key: str,2 max_items: int = 200, max_pages: int = 10, **params):3 """Walk a cursor-paginated Slack method with hard caps on both axes."""4 items, cursor, pages = [], None, 05 while pages < max_pages and len(items) < max_items:6 if cursor:7 params["cursor"] = cursor8 params["limit"] = min(200, max_items - len(items))9 body = await slack_call(method, user_id, **params)10 items.extend(body.get(key, []))11 pages += 112 cursor = body.get("response_metadata", {}).get("next_cursor") or None13 if not cursor:14 break15 return items[:max_items], (cursor is not None)Both caps are load-bearing, and for different reasons. max_pages protects against a server whose cursor never terminates — an infinite loop that burns quota until someone notices the bill. max_items protects the context window: 200 Slack messages at roughly 60 tokens each is about 12,000 tokens, which is a reasonable ceiling for one tool result. Ten thousand messages would be 600,000 tokens, which is three times a large context window and would fail after several minutes of successful API calls.
Return the "there is more" flag to the model as a sentence, not a cursor. "Showing the first 200 of about 4,000 messages; narrow with in:#channel or a date range" is actionable. Handing the model an opaque cursor string invites it to paginate blindly through all 4,000.
Never turn an error into an empty result. An agent cannot tell "nothing matched" from "your token expired", and it will confidently report the first when the truth is the second.
What this means when you build one
Every external-API server you write is really an exercise in deciding what the model is allowed to see, and the default answer should be: much less than the API returns.
A Slack message object has around 30 fields. The model needs four: channel, author, timestamp, text — plus a permalink so the answer can cite a source. Sending the other 26 costs roughly 400 tokens per message instead of 60, which across a 200-message page is 68,000 wasted tokens and a measurably worse answer, because the signal is buried. Normalise aggressively, and treat every field you pass through as something you had to justify.
Three habits follow from that, and they are what separates a server that survives contact with real users from one that dies at 15:00 on a Tuesday:
- Make failure visible. Never convert an error into an empty result. An agent that receives "no results" cannot tell the difference between "nothing matched" and "your token expired", and it will confidently report the first when the truth is the second.
- Put the vendor's constraints in the tool description. The 10-result cap, the search syntax, the fact that only joined channels are searchable — the model can plan around limits it knows about and cannot plan around limits it discovers by failing.
- Instrument the boundary. Log every upstream call with its status, duration and retry count to stderr. When the assistant "feels slow", the answer is almost always one upstream API at the 95th percentile, and without those numbers you are guessing.
The team from the opening changed three things: an 18-per-minute limiter in front of search.messages, full jitter on retries, and a check for ok: false that raised instead of returning an empty list. The 429 rate went to near zero because the requests were spread rather than bunched, and — more importantly — the assistant started saying "Slack is throttling right now, try again in a moment" instead of inventing the absence of an incident.