AI Agent Frameworks

Capstone Project — Building a Multi-Tool Research Agent


The brief is one sentence: build a research agent that answers questions like "How has the UK meal-kit market changed since 2022, and which players are gaining share?" with sourced, current, non-hallucinated answers.

The first version anyone builds takes about forty minutes and mostly works. The version that survives a week of real use takes considerably longer, and the gap between them is entirely made of things that are invisible in a demo: a cache, a rate limiter, a token budget, a structured log, and a deployment target that does not exceed its size limit.

This lesson builds the second version. Every piece of hardening below exists because of a specific way the first version fails.

Five tools behind one sourced answerResearch agentWeb search for current pagesEncyclopedia for backgroundNews search for recent movesPage reader for full textCalculator forthe share maths
Each tool exists to remove one hallucination route — currency, background, recency, detail and arithmetic.

Architecture

Text
                       user question                             |                     [ rate limiter ]  per-user token bucket                             |                       [ cache ]  question hash -> cached answer                             |                     [ agent loop ]  max 6 steps, 45s, 60k token budget                             |        +--------+-----------+-----------+-----------+        |        |           |           |           |    search_web  wiki_lookup  news_search  calculate  read_url        |        |           |           |           |        +--------+-----------+-----------+-----------+                             |                    [ structured logger ]                             |                     answer + sources + cost

Five tools, chosen so that each covers a gap the others leave.

ToolAnswersFails atTypical latency
search_webCurrent, broad, "what is being said about X"Depth; snippets are 200 characters1.5–3 s
wiki_lookupStable background: definitions, history, entitiesAnything from the last few months0.4–1 s
news_searchDated events with publication timestampsAnything older than the index0.5–1.5 s
read_urlThe full text behind a promising search hitPaywalls, JavaScript-rendered pages1–4 s
calculateExact arithmetic on retrieved figuresAnything that is not arithmetic< 1 ms

The overlap between search_web and news_search is deliberate but it needs managing: without descriptions that draw a clear line, the agent calls both for every question and doubles your latency for no gain.

The tools

Bash
pip install langchain langchain-anthropic langgraph ddgs wikipedia-api \            requests beautifulsoup4 python-dotenv

Web search

Python
from langchain_core.tools import toolfrom ddgs import DDGSimport time@tooldef search_web(query: str) -> str:    """Search the live web for current information: market data, company news,    prices, recent events, opinion.    Use for anything that may have changed recently. Do NOT use for    encyclopedic background (use wiki_lookup) or arithmetic (use calculate).    Query should be 3-8 keywords, no question marks.    Returns up to 5 results as: index, title, 200-char snippet, URL.    Follow up with read_url on the most promising result if you need detail.    """    for attempt in range(3):        try:            with DDGS() as d:                hits = list(d.text(query, max_results=5))            break        except Exception:            if attempt == 2:                return "ERROR: web search unavailable. Continue without it."            time.sleep(2 ** attempt)    if not hits:        return (f"No results for {query!r}. Try broader keywords once; if that "                f"fails, state that the information is unavailable.")    return "\n".join(f"[{i}] {h['title']}\n    {h['body'][:200]}\n    {h['href']}"                     for i, h in enumerate(hits, 1))

Encyclopedia lookup

Python
import wikipediaapi_wiki = wikipediaapi.Wikipedia(user_agent="research-agent/1.0", language="en")@tooldef wiki_lookup(topic: str) -> str:    """Look up stable background knowledge: definitions, history, how something    works, who or what an entity is.    Use for context that does not change month to month. Do NOT use for    current prices, recent events or market share figures.    Pass an exact article title if you know it, otherwise a short noun phrase.    Returns the first ~1,200 characters of the summary plus the article URL.    """    page = _wiki.page(topic.strip())    if not page.exists():        return (f"No Wikipedia article for {topic!r}. Try a broader term, or "                f"use search_web instead.")    return f"{page.title}\n{page.summary[:1200]}\n\nSource: {page.fullurl}"

News search

Python
import os, requests@tooldef news_search(query: str, days: int = 30) -> str:    """Search recent news articles with publication dates.    Use when the question involves timing - "when did X happen", "recent    announcements", "this quarter". Prefer this over search_web when a date    matters, because results carry timestamps.    days: look-back window, 1 to 90, default 30.    Returns up to 5 articles as: date, source, headline, URL.    """    days = max(1, min(90, days))    key = os.environ.get("NEWS_API_KEY")    if not key:        return "ERROR: news search not configured. Use search_web instead."    try:        r = requests.get("https://newsapi.org/v2/everything",                         params={"q": query, "pageSize": 5,                                 "sortBy": "publishedAt", "language": "en"},                         headers={"X-Api-Key": key}, timeout=8)        if r.status_code == 429:            return "ERROR: news rate limit hit. Use search_web instead."        r.raise_for_status()    except requests.RequestException as e:        return f"ERROR: news search failed ({type(e).__name__}). Use search_web."    arts = r.json().get("articles", [])    if not arts:        return f"No news for {query!r} in the last {days} days."    return "\n".join(        f"{a['publishedAt'][:10]} | {a['source']['name']} | {a['title']}\n"        f"    {a['url']}" for a in arts)

Page reader

Python
from bs4 import BeautifulSoup@tooldef read_url(url: str) -> str:    """Fetch a web page and return its main text, truncated to 3,000 characters.    Use after search_web when a snippet looks promising but is too short.    Only pass URLs that appeared in an earlier tool result. Never invent a URL.    """    if not url.startswith(("http://", "https://")):        return "ERROR: url must start with http:// or https://"    try:        r = requests.get(url, timeout=10,                         headers={"User-Agent": "research-agent/1.0"})        r.raise_for_status()    except requests.RequestException as e:        return f"ERROR: could not fetch {url} ({type(e).__name__}). Try another."    soup = BeautifulSoup(r.text, "html.parser")    for junk in soup(["script", "style", "nav", "footer", "header", "aside"]):        junk.decompose()    text = " ".join(soup.get_text(" ").split())    if len(text) < 200:        return f"Page at {url} had almost no extractable text (likely JavaScript-rendered)."    return text[:3000] + ("… [truncated]" if len(text) > 3000 else "")

The 3,000-character cap is not arbitrary. A full news article is often 20,000 characters, or roughly 5,000 tokens. Because the agent re-sends its whole scratchpad on every step, one uncapped read_url at step 2 adds 5,000 tokens to steps 3, 4, 5 and 6 as well — 25,000 tokens across the run from one call. The cap turns that into about 3,750.

Calculator

Python
import ast, operator_OPS = {ast.Add: operator.add, ast.Sub: operator.sub, ast.Mult: operator.mul,        ast.Div: operator.truediv, ast.Pow: operator.pow, ast.USub: operator.neg}def _ev(n):    if isinstance(n, ast.Constant) and isinstance(n.value, (int, float)):        return n.value    if isinstance(n, ast.BinOp) and type(n.op) in _OPS:        return _OPS[type(n.op)](_ev(n.left), _ev(n.right))    if isinstance(n, ast.UnaryOp) and type(n.op) in _OPS:        return _OPS[type(n.op)](_ev(n.operand))    raise ValueError("only numbers and + - * / ** are allowed")@tooldef calculate(expression: str) -> str:    """Evaluate an arithmetic expression exactly.    ALWAYS use this for every calculation, including percentages, growth rates    and differences. Never compute a figure yourself.    Write numbers in full: 1450000000, not 1.45B. Percentages as decimals.    Example: '(1450000000 - 1120000000) / 1120000000'.    """    try:        return f"{_ev(ast.parse(expression, mode='eval').body):,.6f}".rstrip("0").rstrip(".")    except Exception as e:        return f"ERROR: {e}. Send digits and + - * / ** ( ) only."

eval() would be four characters shorter and would hand anyone who can influence the model's output a shell on your server. The AST walker accepts arithmetic and nothing else.

Assembling the agent

Python
import osfrom langchain_anthropic import ChatAnthropicfrom langchain.agents import create_agentfrom langchain.agents.middleware import ModelCallLimitMiddlewareSYSTEM = """You are a research analyst. Produce sourced, current answers.Method:1. Use wiki_lookup for background you need but do not have.2. Use search_web or news_search for anything current. Use news_search when   dates matter.3. Use read_url on at most two promising results for depth.4. Use calculate for every number you derive.5. Answer in under 250 words, then list sources as URLs.Hard rules:- Never state a figure that did not come from a tool result.- Never invent a URL. Only read URLs that appeared in a tool result.- If sources disagree, say so and give both figures with their sources.- If you cannot verify something, say "not verified" rather than estimating."""TOOLS = [search_web, wiki_lookup, news_search, read_url, calculate]def build_agent():    llm = ChatAnthropic(model=os.getenv("AGENT_MODEL", "claude-sonnet-5"),                        max_tokens=1500)    return create_agent(        model=llm,        tools=TOOLS,        system_prompt=SYSTEM,        middleware=[ModelCallLimitMiddleware(run_limit=6, exit_behavior="end")],    )def tool_steps(messages):    """(tool name, result text) for every tool call in a run, in order."""    return [(m.name, m.content) for m in messages if m.type == "tool"]

Six model calls, and a 45-second deadline that the caller enforces with asyncio.wait_for around agent.ainvoke(...), because create_agent has no deadline setting of its own. Work out what those bound. Six steps at roughly 4,000 input tokens each averaged over the run is 24,000 input tokens, plus perhaps 3,000 output. At example rates of 3 dollars per million input and 15 dollars per million output (check current prices for your model) that is 0.072+0.045=0.1170.072 + 0.045 = 0.117 dollars — about 12 cents as an absolute worst case per question. Without the caps, the worst case is bounded only by your quota.

Every tool must be able to fail in a sentence the model can act on. A raised exception ends the request; a returned string beginning with ERROR keeps the agent working.

Hardening

Caching

Research questions repeat. In a real deployment, roughly 30 per cent of questions in a day are near-repeats of an earlier one.

Python
import hashlib, json, time, sqlite3class AnswerCache:    def __init__(self, path="cache.db", ttl_seconds=3600):        self.db = sqlite3.connect(path, check_same_thread=False)        self.db.execute("CREATE TABLE IF NOT EXISTS cache "                        "(k TEXT PRIMARY KEY, v TEXT, ts REAL)")        self.ttl = ttl_seconds        self.hits = self.misses = 0    @staticmethod    def _key(question: str) -> str:        norm = " ".join(question.lower().split())        return hashlib.sha256(norm.encode()).hexdigest()    def get(self, question: str):        row = self.db.execute("SELECT v, ts FROM cache WHERE k=?",                              (self._key(question),)).fetchone()        if row and time.time() - row[1] < self.ttl:            self.hits += 1            return json.loads(row[0])        self.misses += 1        return None    def put(self, question: str, answer: dict):        self.db.execute("INSERT OR REPLACE INTO cache VALUES (?,?,?)",                        (self._key(question), json.dumps(answer), time.time()))        self.db.commit()

The TTL is the interesting parameter, and one hour is a compromise you should adjust per question type. A cached answer to "what is a meal kit" is good for a month. A cached answer to "what happened to oil prices today" is stale in an hour. If you can classify the question, set the TTL from the classification — 24 hours for background, 15 minutes for anything with "today", "latest" or "current" in it.

The arithmetic at a 30 per cent hit rate: 10,000 questions a day at 0.06 dollars average without caching is 600 dollars; with caching, 7,000 live questions is 420 dollars — 180 dollars a day, or roughly 65,000 dollars a year, from about 30 lines of code.

Rate limiting

Python
import threadingclass TokenBucket:    """Allows `rate` requests per second with bursts up to `capacity`."""    def __init__(self, rate=2.0, capacity=5):        self.rate, self.capacity = rate, capacity        self.tokens = float(capacity)        self.updated = time.monotonic()        self.lock = threading.Lock()    def take(self, n=1.0) -> bool:        with self.lock:            now = time.monotonic()            self.tokens = min(self.capacity,                              self.tokens + (now - self.updated) * self.rate)            self.updated = now            if self.tokens >= n:                self.tokens -= n                return True            return False_buckets = {}def allow(user_id: str) -> bool:    return _buckets.setdefault(user_id, TokenBucket()).take()

Two rates per second with a burst of five means a user can fire five questions immediately, then one every half second. That absorbs a genuine flurry of interest while stopping a script from issuing 500 questions a minute. Rate-limit outbound calls too — the news API will start returning 429 long before your own limits bite.

Structured logging

Python
import logging, uuid, jsonlog = logging.getLogger("agent")def answer(question: str, user_id: str, cache, agent) -> dict:    rid = uuid.uuid4().hex[:12]    t0 = time.perf_counter()    if not allow(user_id):        log.warning(json.dumps({"rid": rid, "event": "rate_limited",                                "user": user_id}))        return {"error": "Too many requests. Try again in a few seconds."}    cached = cache.get(question)    if cached:        log.info(json.dumps({"rid": rid, "event": "cache_hit",                             "ms": round((time.perf_counter() - t0) * 1000)}))        return cached    try:        out = agent.invoke({"messages": [{"role": "user", "content": question}]})    except Exception as e:        log.error(json.dumps({"rid": rid, "event": "agent_error",                              "type": type(e).__name__, "msg": str(e)[:200]}))        return {"error": "The research agent failed. Please retry."}    msgs = out["messages"]    model_turns = sum(m.type == "ai" for m in msgs)    result = {        "answer": msgs[-1].content,        "tools_used": [name for name, _ in tool_steps(msgs)],        "steps": model_turns,        "hit_cap": model_turns >= 6,    }    log.info(json.dumps({"rid": rid, "event": "answered",                         "ms": round((time.perf_counter() - t0) * 1000),                         **{k: v for k, v in result.items() if k != "answer"}}))    cache.put(question, result)    return result

One JSON object per line, one request id threading through everything. This is what makes a support ticket answerable: given the request id, you can see every tool the agent called, in order, with timings. hit_cap is worth logging explicitly — a rising proportion of requests hitting the six-call ceiling is the earliest signal that question complexity has outgrown your budget.

Caps, caches and rate limits are invisible in a demo and are the whole difference between an agent that costs pennies and one that costs hundreds.

Deployment

Serverless functionContainer service
Cold start2–8 s with a heavy dependency treeNone once warm
Size limitOften ~250 MB unzipped — a real constraintEffectively none
Max durationCommonly 10–60 s — your 45 s cap must fit inside itYour choice
Cache and rate-limit stateMust be external (Redis); local files do not persistLocal works, external is better
Cost at low trafficNear zeroYou pay for idle
Cost at steady trafficHigher per requestLower

Three traps catch people specifically with agents.

Timeouts must nest. If the platform kills the function at 30 seconds and your agent's deadline is 45, the platform wins and you get a truncated request with no log line explaining why. Set the agent's cap below the platform's — 25 against 30 — so your code owns the failure and can log it.

Local state does not exist. An SQLite cache and an in-process token bucket both live on one instance. With ten concurrent instances you get ten independent caches (hit rate collapses) and ten independent rate limiters (a user can send ten times your intended rate). Move both to Redis before you scale out.

Size limits bite late. The full agent stack with HTML parsing and HTTP libraries can approach or exceed a 250 MB unzipped limit once transitive dependencies land. Check du -sh on your installed packages on day one, not in week six.

Python
# api/research.py - serverless entry pointfrom http.server import BaseHTTPRequestHandlerimport json_agent = build_agent()_cache = AnswerCache(path="/tmp/cache.db")   # per-instance; use Redis in productionclass handler(BaseHTTPRequestHandler):    def do_POST(self):        body = json.loads(self.rfile.read(            int(self.headers.get("Content-Length", 0))) or b"{}")        q = (body.get("question") or "").strip()        if not q or len(q) > 500:            self._send(400, {"error": "question must be 1-500 characters"})            return        self._send(200, answer(q, body.get("user_id", "anon"), _cache, _agent))    def _send(self, code, payload):        data = json.dumps(payload).encode()        self.send_response(code)        self.send_header("Content-Type", "application/json")        self.send_header("Content-Length", str(len(data)))        self.end_headers()        self.wfile.write(data)

Note the 500-character input cap. Without it, someone pastes a 40,000-character document into question and every one of your six agent steps carries it. That is one request costing several dollars, and it is trivially avoidable.

Knowing whether it actually works

The agent will look impressive on the questions you thought of while building it. That tells you nothing, because you unconsciously chose questions it handles. Build a fixed set of fifteen questions covering the cases you did not design for: a question with no answer, a question whose answer changed last week, a question where two sources disagree, a question needing arithmetic on two retrieved figures, and a question whose obvious search terms return junk.

Record four things per question: which tools ran, whether every figure in the answer appears in some tool result, whether the cited URLs appeared in tool results, and the latency. The second and third are the ones that matter — they detect hallucination mechanically, without a human reading the answer, and they are checkable with a substring test:

Python
def unsupported_numbers(answer_text: str, steps) -> list[str]:    import re    observations = " ".join(str(s[1]) for s in steps)    figures = re.findall(r"\d[\d,]*\.?\d*", answer_text)    return [f for f in figures if f.replace(",", "") not in            observations.replace(",", "")]

Any figure this returns is a figure the agent produced without evidence. On a well-built agent it returns an empty list almost always; when it does not, you have caught the exact failure that makes research agents dangerous — a confident, sourced-looking answer with an invented number in it. Run this on every response in production, log the failures, and you will know about the problem before a user does.