Model Context Protocol (MCP)

Integrating External APIs (Google, Slack, and Beyond)


At 14:50 the Slack MCP server is fine. At 15:00, when the second standup ends and four analysts start asking the assistant questions at once, every slack_search call starts coming back empty. No exception, no error in the agent's transcript — the tool returns "No messages found" and the model confidently tells three people that the incident was never discussed.

Two things went wrong, and both are typical. First, the arithmetic: search.messages sits in a Slack rate-limit tier allowing roughly 20 requests per minute per workspace. Each analyst question fans out to eight searches. Four analysts asking one question a minute is 32 requests a minute against a ceiling of 20, so twelve requests are rejected every minute with HTTP 429 and a Retry-After: 30 header.

Second, and worse: the server retried those twelve after exactly 30 seconds — all twelve at the same instant, because they all failed at the same instant and all waited the same amount. The retry burst collided with the next minute's fresh traffic and the failure rate went up, not down. Meanwhile the handler treated a non-200 response as "no results" and returned a cheerful empty list, so nothing anywhere in the system said the word "failing".

An MCP server that wraps a third-party API is not a thin proxy. It is the component that turns an unreliable, rate-limited, paginated, idiosyncratic HTTP API into something a language model can use without lying to its user.

Why the assistant said "no messages found"Four analysts callslack_search at onceSlackreplies 429,not an exceptionServer swallowsit, returns emptyModelreports theabsence as factFix: limit by construction, retry with backoff and jitter, and return an error the model can see.
An empty result and a rate-limited result must never look the same to the model — one is data, the other is a failure.

What actually belongs in the server

LayerJobWhat breaks without it
Credential managementHold keys and refresh tokens; never surface them as tool argumentsThe model is asked to supply a secret, and eventually invents one
Client-side rate limitingStay under the quota before sendingYou discover the limit by being throttled
Retry with backoff and jitterRecover from transient failures without synchronised burstsRetry storms turn a blip into an outage
Error mappingTurn HTTP and vendor error codes into sentences the model can act on"Error 400" produces an identical retry, forever
PaginationWalk cursors with a hard page capOne call fetches 40,000 messages into the context window
NormalisationStrip vendor payloads to the fields that matter90% of the tokens are metadata the model ignores

The model should never see an HTTP status code, a cursor token, or a vendor-specific error string. Those are the server's problem, and hiding them is most of its value.

Authentication patterns

API keys: for APIs that identify an application

Google Custom Search, most weather APIs, most search APIs. One key, loaded from the environment, applied to every request. The rule that matters: the key never appears in the tool's input schema. If api_key is a parameter, the model will eventually be asked to fill it in, and a model asked for a credential will hallucinate a plausible-looking string.

Python
import os, httpxGOOGLE_KEY = os.environ["GOOGLE_API_KEY"]     # fail fast at startup, not mid-callGOOGLE_CX  = os.environ["GOOGLE_CSE_ID"]client = httpx.AsyncClient(    timeout=httpx.Timeout(connect=3.0, read=15.0, write=5.0, pool=5.0),    limits=httpx.Limits(max_connections=20, max_keepalive_connections=10),)

Note the split timeouts. A single timeout=30 means a dead host ties up a worker for thirty seconds; a three-second connect timeout fails fast on unreachable hosts while still allowing a slow search fifteen seconds to think.

OAuth 2.0: for APIs that act on behalf of a person

Slack, Google Drive, GitHub, Notion. The server holds a long-lived refresh token and exchanges it for short-lived access tokens. The part people get wrong is refreshing reactively — waiting for a 401 and then refreshing — which means every session pays at least one failed request, and concurrent calls all try to refresh at once.

Python
import asyncio, timeclass TokenStore:    """One access token per user, refreshed proactively under a lock."""    def __init__(self):        self._tokens: dict[str, tuple[str, float]] = {}   # user -> (token, expires_at)        self._locks: dict[str, asyncio.Lock] = {}    async def get(self, user_id: str) -> str:        token, expires_at = self._tokens.get(user_id, (None, 0.0))        # Refresh 60 s early so an in-flight request never expires mid-call.        if token and time.time() < expires_at - 60:            return token        lock = self._locks.setdefault(user_id, asyncio.Lock())        async with lock:            # Re-check: another coroutine may have refreshed while we waited.            token, expires_at = self._tokens.get(user_id, (None, 0.0))            if token and time.time() < expires_at - 60:                return token            token, ttl = await self._refresh(user_id)            self._tokens[user_id] = (token, time.time() + ttl)            return token    async def _refresh(self, user_id: str) -> tuple[str, float]:        r = await client.post("https://slack.com/api/oauth.v2.access", data={            "grant_type": "refresh_token",            "refresh_token": await load_refresh_token(user_id),   # from encrypted store            "client_id": os.environ["SLACK_CLIENT_ID"],            "client_secret": os.environ["SLACK_CLIENT_SECRET"],        })        body = r.json()        if not body.get("ok"):            raise AuthError(f"token refresh failed: {body.get('error')}")        await save_refresh_token(user_id, body["refresh_token"])   # rotation        return body["access_token"], float(body.get("expires_in", 43200))

Three details earn their place. The 60-second early refresh removes the race where a token expires between the check and the request arriving at the server. The double-checked lock stops ten concurrent tool calls from firing ten refreshes — which with rotating refresh tokens invalidates nine of them and locks the user out. And save_refresh_token runs on every refresh because many providers rotate the refresh token too; ignore the new one and the integration silently dies the next time the old one is used.

For a multi-user server, the identity comes from the access token that arrives with each MCP request, never from a tool argument. A tool called read_slack_as(user_id) is an impersonation API with extra steps.

Rate limits and retries are one problem

People implement these separately and then discover they fight each other. A retry policy without a rate limiter reacts to throttling by generating more traffic. A rate limiter without retries drops legitimate work on the first transient blip.

The limiter: stay under the ceiling by construction

Python
import asyncio, timeclass AsyncRateLimiter:    """Token bucket. Callers await capacity instead of being rejected."""    def __init__(self, rate_per_sec: float, burst: float):        self.rate, self.burst = rate_per_sec, burst        self.tokens, self.updated = burst, time.monotonic()        self._lock = asyncio.Lock()    async def acquire(self, cost: float = 1.0) -> None:        async with self._lock:            while True:                now = time.monotonic()                self.tokens = min(self.burst,                                  self.tokens + (now - self.updated) * self.rate)                self.updated = now                if self.tokens >= cost:                    self.tokens -= cost                    return                await asyncio.sleep((cost - self.tokens) / self.rate)# Slack tier-2 methods: ~20 per minute. Leave headroom.slack_search_limiter = AsyncRateLimiter(rate_per_sec=18 / 60, burst=5)

Set the rate slightly below the documented ceiling — 18 per minute against a limit of 20 — because your clock and the vendor's window boundaries do not line up, and running at exactly the limit guarantees occasional 429s.

The retry: back off, and add jitter

Exponential backoff doubles the wait after each failure: 1 s, 2 s, 4 s, 8 s, 16 s. Five attempts span about 31 seconds of waiting. That is the right shape, and on its own it is what caused the outage in the opening — twelve simultaneous failures produce twelve simultaneous retries.

Full jitter fixes it by waiting a uniformly random time between zero and the exponential ceiling. The expected wait is halved, and twelve retries spread themselves across the whole window instead of stacking on one instant.

AttemptFixed 30 sExponential, no jitterFull jitter: random(0, 2^n)
130.0 s1.0 suniform in [0, 1] s
230.0 s2.0 suniform in [0, 2] s
330.0 s4.0 suniform in [0, 4] s
430.0 s8.0 suniform in [0, 8] s
530.0 s16.0 suniform in [0, 16] s
Collision behaviourAll retries land togetherAll retries land togetherSpread across the window
Python
import randomfrom tenacity import (retry, stop_after_attempt, wait_random_exponential,                      retry_if_exception_type)class Retryable(Exception):    """Transient. Worth another attempt."""    def __init__(self, msg, retry_after: float | None = None):        super().__init__(msg)        self.retry_after = retry_afterclass Permanent(Exception):    """Retrying will produce the identical failure."""@retry(retry=retry_if_exception_type(Retryable),       wait=wait_random_exponential(multiplier=1, max=16),   # full jitter       stop=stop_after_attempt(5),       reraise=True)async def call_api(method: str, url: str, **kw) -> dict:    r = await client.request(method, url, **kw)    if r.status_code == 429:        # Always prefer the server's own advice, then add jitter on top.        wait = float(r.headers.get("Retry-After", 5))        await asyncio.sleep(wait + random.uniform(0, 1.0))        raise Retryable(f"rate limited, server asked for {wait}s", retry_after=wait)    if 500 <= r.status_code < 600:        raise Retryable(f"upstream {r.status_code}")    if r.status_code in (401, 403):        raise Permanent(f"not authorised ({r.status_code}) - check scopes: {r.text[:200]}")    if 400 <= r.status_code < 500:        raise Permanent(f"request rejected ({r.status_code}): {r.text[:200]}")    return r.json()

Which failures are worth retrying

ResponseRetry?WhyWhat the model should be told
429 Too Many RequestsYes, after Retry-AfterPurely temporalOnly after retries are exhausted: "busy, try again shortly"
500 / 502 / 503 / 504Yes, with backoffUpstream instability"The service is unavailable right now"
Connection reset / timeoutYes, if the call is idempotentNetwork, not logicNothing, if a retry succeeds
401 UnauthorizedOnce, after refreshing the tokenUsually an expired token"Reconnect your account"
403 ForbiddenNoA missing scope will still be missing"This needs the files:read scope"
404 Not FoundNoThe object does not exist"No channel named #incidnets — did you mean #incidents?"
400 Bad RequestNoThe arguments are wrongThe vendor's validation message, verbatim

The distinction is the whole point of the two exception classes. Retrying a 403 five times with backoff wastes 31 seconds to arrive at the same 403, and the agent's user watches a spinner for half a minute to learn nothing.

A rate limiter without retries drops legitimate work; retries without a rate limiter answer throttling by sending more traffic. Ship them together or ship neither.

A Google search server

Google's Custom Search JSON API returns at most 10 results per request and at most 100 in total, and the free tier is 100 queries per day. Those constraints belong in the tool description, because a model that does not know about them will paginate happily into a wall.

Python
from mcp.server.mcpserver import MCPServermcp = MCPServer("web-search")google_limiter = AsyncRateLimiter(rate_per_sec=1.0, burst=3)@mcp.tool()async def web_search(query: str, num_results: int = 5,                     site: str | None = None) -> str:    """Search the public web and return titles, URLs and snippets.    Returns at most 10 results per call, newest ranking first.    Use `site` to restrict to one domain, e.g. site="docs.python.org".    Snippets are ~150 characters; fetch the URL for full text.    """    num_results = max(1, min(num_results, 10))    q = f"site:{site} {query}" if site else query    await google_limiter.acquire()    try:        data = await call_api("GET", "https://www.googleapis.com/customsearch/v1",                              params={"key": GOOGLE_KEY, "cx": GOOGLE_CX,                                      "q": q, "num": num_results})    except Permanent as e:        if "quota" in str(e).lower():               # Google reports quota as a 403            return ("The daily search quota is exhausted. Answer from what you "                    "already have, or ask the user to try tomorrow.")        return f"Search failed and retrying will not help: {e}"    except Retryable as e:        return f"Search is temporarily unavailable ({e}). Try again in a minute."    items = data.get("items", [])    if not items:        total = data.get("searchInformation", {}).get("totalResults", "0")        return (f"No results for {q!r} (Google reported {total} total). "                f"Try broader terms or drop the site restriction.")    return "\n\n".join(        f"{i}. {it['title']}\n   {it['link']}\n   {it.get('snippet','').strip()}"        for i, it in enumerate(items, 1))

The quota message is the interesting line. "Quota exceeded" tells the model nothing useful; "answer from what you already have, or ask the user to try tomorrow" tells it what to do instead of retrying.

A Slack server, and Slack's favourite trap

Slack returns HTTP 200 with {"ok": false, "error": "..."} for most application-level failures. A handler that checks only r.status_code == 200 treats "channel_not_found" and "missing_scope" as success and returns an empty result — which is precisely the bug that let the agent tell three people the incident was never discussed.

Python
SLACK_ERRORS = {    "channel_not_found": "That channel does not exist or the app is not a member of it.",    "not_in_channel":    "The app must be invited to that channel first (/invite @bot).",    "missing_scope":     "The Slack app is missing a required OAuth scope.",    "invalid_auth":      "The Slack token is invalid or revoked. Reconnect the workspace.",    "ratelimited":       "Slack is throttling this workspace. Wait a moment.",}async def slack_call(method: str, user_id: str, **params) -> dict:    await slack_search_limiter.acquire()    token = await tokens.get(user_id)    body = await call_api("POST", f"https://slack.com/api/{method}",                          headers={"Authorization": f"Bearer {token}"},                          data=params)    if not body.get("ok"):        err = body.get("error", "unknown_error")        raise Permanent(SLACK_ERRORS.get(err, f"Slack rejected the call: {err}"))    return bodyfrom mcp.server.auth.middleware.auth_context import get_access_token@mcp.tool()async def slack_search(query: str, count: int = 10) -> str:    """Search messages across channels the app can see.    Supports Slack search syntax: in:#channel, from:@user, before:2026-08-01.    Returns at most 20 messages, most relevant first, with permalinks.    """    count = max(1, min(count, 20))    access = get_access_token()        # the caller's validated OAuth token, per request    if access is None or access.subject is None:        return "This tool needs a signed-in user. Ask the user to connect their account."    try:        body = await slack_call("search.messages", access.subject,                                query=query, count=count, sort="score")    except Permanent as e:        return f"Slack search failed: {e}"    matches = body["messages"]["matches"]    total = body["messages"]["total"]    if not matches:        return (f"No messages matched {query!r}. Slack search only covers channels "                f"the app has joined - try naming a channel with in:#name.")    lines = [f"{total} match(es); showing {len(matches)}:"]    for m in matches:        text = " ".join(m.get("text", "").split())[:400]        lines.append(f"[#{m['channel']['name']}  {m['username']}  "                     f"{m['ts'][:10]}]\n{text}\n{m['permalink']}")    return "\n\n".join(lines)

The user's identity comes from get_access_token(), which the SDK fills in from the bearer token it validated for this request; the model never sees or supplies it. The empty-result message does real work: it explains why a search might legitimately find nothing and suggests a narrower query. That single sentence converts a dead end into a next step.

Pagination

Three schemes dominate, and each fails differently.

SchemeLooks likeUsed byFailure mode
Cursornext_cursor in the responseSlack, StripeCursors expire; resuming an old one 400s
Offset / pagestart=11, page=3Google CSE, older REST APIsRows inserted mid-walk cause duplicates and skips
Link headerLink: <...>; rel="next"GitHubEasy to parse wrongly; the URL must be used verbatim
Python
async def paginate_slack(method: str, user_id: str, key: str,                         max_items: int = 200, max_pages: int = 10, **params):    """Walk a cursor-paginated Slack method with hard caps on both axes."""    items, cursor, pages = [], None, 0    while pages < max_pages and len(items) < max_items:        if cursor:            params["cursor"] = cursor        params["limit"] = min(200, max_items - len(items))        body = await slack_call(method, user_id, **params)        items.extend(body.get(key, []))        pages += 1        cursor = body.get("response_metadata", {}).get("next_cursor") or None        if not cursor:            break    return items[:max_items], (cursor is not None)

Both caps are load-bearing, and for different reasons. max_pages protects against a server whose cursor never terminates — an infinite loop that burns quota until someone notices the bill. max_items protects the context window: 200 Slack messages at roughly 60 tokens each is about 12,000 tokens, which is a reasonable ceiling for one tool result. Ten thousand messages would be 600,000 tokens, which is three times a large context window and would fail after several minutes of successful API calls.

Return the "there is more" flag to the model as a sentence, not a cursor. "Showing the first 200 of about 4,000 messages; narrow with in:#channel or a date range" is actionable. Handing the model an opaque cursor string invites it to paginate blindly through all 4,000.

Never turn an error into an empty result. An agent cannot tell "nothing matched" from "your token expired", and it will confidently report the first when the truth is the second.

What this means when you build one

Every external-API server you write is really an exercise in deciding what the model is allowed to see, and the default answer should be: much less than the API returns.

A Slack message object has around 30 fields. The model needs four: channel, author, timestamp, text — plus a permalink so the answer can cite a source. Sending the other 26 costs roughly 400 tokens per message instead of 60, which across a 200-message page is 68,000 wasted tokens and a measurably worse answer, because the signal is buried. Normalise aggressively, and treat every field you pass through as something you had to justify.

Three habits follow from that, and they are what separates a server that survives contact with real users from one that dies at 15:00 on a Tuesday:

  • Make failure visible. Never convert an error into an empty result. An agent that receives "no results" cannot tell the difference between "nothing matched" and "your token expired", and it will confidently report the first when the truth is the second.
  • Put the vendor's constraints in the tool description. The 10-result cap, the search syntax, the fact that only joined channels are searchable — the model can plan around limits it knows about and cannot plan around limits it discovers by failing.
  • Instrument the boundary. Log every upstream call with its status, duration and retry count to stderr. When the assistant "feels slow", the answer is almost always one upstream API at the 95th percentile, and without those numbers you are guessing.

The team from the opening changed three things: an 18-per-minute limiter in front of search.messages, full jitter on retries, and a check for ok: false that raised instead of returning an empty list. The 429 rate went to near zero because the requests were spread rather than bunched, and — more importantly — the assistant started saying "Slack is throttling right now, try again in a moment" instead of inventing the absence of an incident.