Course Content
AI Agent Frameworks
4 sections · 15 lessons
Capstone Project — Building a Multi-Tool Research Agent
The brief is one sentence: build a research agent that answers questions like "How has the UK meal-kit market changed since 2022, and which players are gaining share?" with sourced, current, non-hallucinated answers.
The first version anyone builds takes about forty minutes and mostly works. The version that survives a week of real use takes considerably longer, and the gap between them is entirely made of things that are invisible in a demo: a cache, a rate limiter, a token budget, a structured log, and a deployment target that does not exceed its size limit.
This lesson builds the second version. Every piece of hardening below exists because of a specific way the first version fails.
Architecture
user question | [ rate limiter ] per-user token bucket | [ cache ] question hash -> cached answer | [ agent loop ] max 6 steps, 45s, 60k token budget | +--------+-----------+-----------+-----------+ | | | | | search_web wiki_lookup news_search calculate read_url | | | | | +--------+-----------+-----------+-----------+ | [ structured logger ] | answer + sources + costFive tools, chosen so that each covers a gap the others leave.
| Tool | Answers | Fails at | Typical latency |
|---|---|---|---|
search_web | Current, broad, "what is being said about X" | Depth; snippets are 200 characters | 1.5–3 s |
wiki_lookup | Stable background: definitions, history, entities | Anything from the last few months | 0.4–1 s |
news_search | Dated events with publication timestamps | Anything older than the index | 0.5–1.5 s |
read_url | The full text behind a promising search hit | Paywalls, JavaScript-rendered pages | 1–4 s |
calculate | Exact arithmetic on retrieved figures | Anything that is not arithmetic | < 1 ms |
The overlap between search_web and news_search is deliberate but it needs managing: without descriptions that draw a clear line, the agent calls both for every question and doubles your latency for no gain.
The tools
pip install langchain langchain-anthropic langgraph ddgs wikipedia-api \ requests beautifulsoup4 python-dotenvWeb search
1from langchain_core.tools import tool2from ddgs import DDGS3import time45@tool6def search_web(query: str) -> str:7 """Search the live web for current information: market data, company news,8 prices, recent events, opinion.910 Use for anything that may have changed recently. Do NOT use for11 encyclopedic background (use wiki_lookup) or arithmetic (use calculate).1213 Query should be 3-8 keywords, no question marks.14 Returns up to 5 results as: index, title, 200-char snippet, URL.15 Follow up with read_url on the most promising result if you need detail.16 """17 for attempt in range(3):18 try:19 with DDGS() as d:20 hits = list(d.text(query, max_results=5))21 break22 except Exception:23 if attempt == 2:24 return "ERROR: web search unavailable. Continue without it."25 time.sleep(2 ** attempt)26 if not hits:27 return (f"No results for {query!r}. Try broader keywords once; if that "28 f"fails, state that the information is unavailable.")29 return "\n".join(f"[{i}] {h['title']}\n {h['body'][:200]}\n {h['href']}"30 for i, h in enumerate(hits, 1))Encyclopedia lookup
1import wikipediaapi23_wiki = wikipediaapi.Wikipedia(user_agent="research-agent/1.0", language="en")45@tool6def wiki_lookup(topic: str) -> str:7 """Look up stable background knowledge: definitions, history, how something8 works, who or what an entity is.910 Use for context that does not change month to month. Do NOT use for11 current prices, recent events or market share figures.1213 Pass an exact article title if you know it, otherwise a short noun phrase.14 Returns the first ~1,200 characters of the summary plus the article URL.15 """16 page = _wiki.page(topic.strip())17 if not page.exists():18 return (f"No Wikipedia article for {topic!r}. Try a broader term, or "19 f"use search_web instead.")20 return f"{page.title}\n{page.summary[:1200]}\n\nSource: {page.fullurl}"News search
1import os, requests23@tool4def news_search(query: str, days: int = 30) -> str:5 """Search recent news articles with publication dates.67 Use when the question involves timing - "when did X happen", "recent8 announcements", "this quarter". Prefer this over search_web when a date9 matters, because results carry timestamps.1011 days: look-back window, 1 to 90, default 30.12 Returns up to 5 articles as: date, source, headline, URL.13 """14 days = max(1, min(90, days))15 key = os.environ.get("NEWS_API_KEY")16 if not key:17 return "ERROR: news search not configured. Use search_web instead."18 try:19 r = requests.get("https://newsapi.org/v2/everything",20 params={"q": query, "pageSize": 5,21 "sortBy": "publishedAt", "language": "en"},22 headers={"X-Api-Key": key}, timeout=8)23 if r.status_code == 429:24 return "ERROR: news rate limit hit. Use search_web instead."25 r.raise_for_status()26 except requests.RequestException as e:27 return f"ERROR: news search failed ({type(e).__name__}). Use search_web."28 arts = r.json().get("articles", [])29 if not arts:30 return f"No news for {query!r} in the last {days} days."31 return "\n".join(32 f"{a['publishedAt'][:10]} | {a['source']['name']} | {a['title']}\n"33 f" {a['url']}" for a in arts)Page reader
1from bs4 import BeautifulSoup23@tool4def read_url(url: str) -> str:5 """Fetch a web page and return its main text, truncated to 3,000 characters.67 Use after search_web when a snippet looks promising but is too short.8 Only pass URLs that appeared in an earlier tool result. Never invent a URL.9 """10 if not url.startswith(("http://", "https://")):11 return "ERROR: url must start with http:// or https://"12 try:13 r = requests.get(url, timeout=10,14 headers={"User-Agent": "research-agent/1.0"})15 r.raise_for_status()16 except requests.RequestException as e:17 return f"ERROR: could not fetch {url} ({type(e).__name__}). Try another."18 soup = BeautifulSoup(r.text, "html.parser")19 for junk in soup(["script", "style", "nav", "footer", "header", "aside"]):20 junk.decompose()21 text = " ".join(soup.get_text(" ").split())22 if len(text) < 200:23 return f"Page at {url} had almost no extractable text (likely JavaScript-rendered)."24 return text[:3000] + ("… [truncated]" if len(text) > 3000 else "")The 3,000-character cap is not arbitrary. A full news article is often 20,000 characters, or roughly 5,000 tokens. Because the agent re-sends its whole scratchpad on every step, one uncapped read_url at step 2 adds 5,000 tokens to steps 3, 4, 5 and 6 as well — 25,000 tokens across the run from one call. The cap turns that into about 3,750.
Calculator
1import ast, operator23_OPS = {ast.Add: operator.add, ast.Sub: operator.sub, ast.Mult: operator.mul,4 ast.Div: operator.truediv, ast.Pow: operator.pow, ast.USub: operator.neg}56def _ev(n):7 if isinstance(n, ast.Constant) and isinstance(n.value, (int, float)):8 return n.value9 if isinstance(n, ast.BinOp) and type(n.op) in _OPS:10 return _OPS[type(n.op)](_ev(n.left), _ev(n.right))11 if isinstance(n, ast.UnaryOp) and type(n.op) in _OPS:12 return _OPS[type(n.op)](_ev(n.operand))13 raise ValueError("only numbers and + - * / ** are allowed")1415@tool16def calculate(expression: str) -> str:17 """Evaluate an arithmetic expression exactly.1819 ALWAYS use this for every calculation, including percentages, growth rates20 and differences. Never compute a figure yourself.21 Write numbers in full: 1450000000, not 1.45B. Percentages as decimals.22 Example: '(1450000000 - 1120000000) / 1120000000'.23 """24 try:25 return f"{_ev(ast.parse(expression, mode='eval').body):,.6f}".rstrip("0").rstrip(".")26 except Exception as e:27 return f"ERROR: {e}. Send digits and + - * / ** ( ) only."eval() would be four characters shorter and would hand anyone who can influence the model's output a shell on your server. The AST walker accepts arithmetic and nothing else.
Assembling the agent
1import os2from langchain_anthropic import ChatAnthropic3from langchain.agents import create_agent4from langchain.agents.middleware import ModelCallLimitMiddleware56SYSTEM = """You are a research analyst. Produce sourced, current answers.78Method:91. Use wiki_lookup for background you need but do not have.102. Use search_web or news_search for anything current. Use news_search when11 dates matter.123. Use read_url on at most two promising results for depth.134. Use calculate for every number you derive.145. Answer in under 250 words, then list sources as URLs.1516Hard rules:17- Never state a figure that did not come from a tool result.18- Never invent a URL. Only read URLs that appeared in a tool result.19- If sources disagree, say so and give both figures with their sources.20- If you cannot verify something, say "not verified" rather than estimating."""2122TOOLS = [search_web, wiki_lookup, news_search, read_url, calculate]2324def build_agent():25 llm = ChatAnthropic(model=os.getenv("AGENT_MODEL", "claude-sonnet-5"),26 max_tokens=1500)27 return create_agent(28 model=llm,29 tools=TOOLS,30 system_prompt=SYSTEM,31 middleware=[ModelCallLimitMiddleware(run_limit=6, exit_behavior="end")],32 )3334def tool_steps(messages):35 """(tool name, result text) for every tool call in a run, in order."""36 return [(m.name, m.content) for m in messages if m.type == "tool"]Six model calls, and a 45-second deadline that the caller enforces with asyncio.wait_for around agent.ainvoke(...), because create_agent has no deadline setting of its own. Work out what those bound. Six steps at roughly 4,000 input tokens each averaged over the run is 24,000 input tokens, plus perhaps 3,000 output. At example rates of 3 dollars per million input and 15 dollars per million output (check current prices for your model) that is 0.072+0.045=0.117 dollars — about 12 cents as an absolute worst case per question. Without the caps, the worst case is bounded only by your quota.
Every tool must be able to fail in a sentence the model can act on. A raised exception ends the request; a returned string beginning with ERROR keeps the agent working.
Hardening
Caching
Research questions repeat. In a real deployment, roughly 30 per cent of questions in a day are near-repeats of an earlier one.
1import hashlib, json, time, sqlite323class AnswerCache:4 def __init__(self, path="cache.db", ttl_seconds=3600):5 self.db = sqlite3.connect(path, check_same_thread=False)6 self.db.execute("CREATE TABLE IF NOT EXISTS cache "7 "(k TEXT PRIMARY KEY, v TEXT, ts REAL)")8 self.ttl = ttl_seconds9 self.hits = self.misses = 01011 @staticmethod12 def _key(question: str) -> str:13 norm = " ".join(question.lower().split())14 return hashlib.sha256(norm.encode()).hexdigest()1516 def get(self, question: str):17 row = self.db.execute("SELECT v, ts FROM cache WHERE k=?",18 (self._key(question),)).fetchone()19 if row and time.time() - row[1] < self.ttl:20 self.hits += 121 return json.loads(row[0])22 self.misses += 123 return None2425 def put(self, question: str, answer: dict):26 self.db.execute("INSERT OR REPLACE INTO cache VALUES (?,?,?)",27 (self._key(question), json.dumps(answer), time.time()))28 self.db.commit()The TTL is the interesting parameter, and one hour is a compromise you should adjust per question type. A cached answer to "what is a meal kit" is good for a month. A cached answer to "what happened to oil prices today" is stale in an hour. If you can classify the question, set the TTL from the classification — 24 hours for background, 15 minutes for anything with "today", "latest" or "current" in it.
The arithmetic at a 30 per cent hit rate: 10,000 questions a day at 0.06 dollars average without caching is 600 dollars; with caching, 7,000 live questions is 420 dollars — 180 dollars a day, or roughly 65,000 dollars a year, from about 30 lines of code.
Rate limiting
1import threading23class TokenBucket:4 """Allows `rate` requests per second with bursts up to `capacity`."""5 def __init__(self, rate=2.0, capacity=5):6 self.rate, self.capacity = rate, capacity7 self.tokens = float(capacity)8 self.updated = time.monotonic()9 self.lock = threading.Lock()1011 def take(self, n=1.0) -> bool:12 with self.lock:13 now = time.monotonic()14 self.tokens = min(self.capacity,15 self.tokens + (now - self.updated) * self.rate)16 self.updated = now17 if self.tokens >= n:18 self.tokens -= n19 return True20 return False2122_buckets = {}23def allow(user_id: str) -> bool:24 return _buckets.setdefault(user_id, TokenBucket()).take()Two rates per second with a burst of five means a user can fire five questions immediately, then one every half second. That absorbs a genuine flurry of interest while stopping a script from issuing 500 questions a minute. Rate-limit outbound calls too — the news API will start returning 429 long before your own limits bite.
Structured logging
1import logging, uuid, json23log = logging.getLogger("agent")45def answer(question: str, user_id: str, cache, agent) -> dict:6 rid = uuid.uuid4().hex[:12]7 t0 = time.perf_counter()89 if not allow(user_id):10 log.warning(json.dumps({"rid": rid, "event": "rate_limited",11 "user": user_id}))12 return {"error": "Too many requests. Try again in a few seconds."}1314 cached = cache.get(question)15 if cached:16 log.info(json.dumps({"rid": rid, "event": "cache_hit",17 "ms": round((time.perf_counter() - t0) * 1000)}))18 return cached1920 try:21 out = agent.invoke({"messages": [{"role": "user", "content": question}]})22 except Exception as e:23 log.error(json.dumps({"rid": rid, "event": "agent_error",24 "type": type(e).__name__, "msg": str(e)[:200]}))25 return {"error": "The research agent failed. Please retry."}2627 msgs = out["messages"]28 model_turns = sum(m.type == "ai" for m in msgs)29 result = {30 "answer": msgs[-1].content,31 "tools_used": [name for name, _ in tool_steps(msgs)],32 "steps": model_turns,33 "hit_cap": model_turns >= 6,34 }35 log.info(json.dumps({"rid": rid, "event": "answered",36 "ms": round((time.perf_counter() - t0) * 1000),37 **{k: v for k, v in result.items() if k != "answer"}}))38 cache.put(question, result)39 return resultOne JSON object per line, one request id threading through everything. This is what makes a support ticket answerable: given the request id, you can see every tool the agent called, in order, with timings. hit_cap is worth logging explicitly — a rising proportion of requests hitting the six-call ceiling is the earliest signal that question complexity has outgrown your budget.
Caps, caches and rate limits are invisible in a demo and are the whole difference between an agent that costs pennies and one that costs hundreds.
Deployment
| Serverless function | Container service | |
|---|---|---|
| Cold start | 2–8 s with a heavy dependency tree | None once warm |
| Size limit | Often ~250 MB unzipped — a real constraint | Effectively none |
| Max duration | Commonly 10–60 s — your 45 s cap must fit inside it | Your choice |
| Cache and rate-limit state | Must be external (Redis); local files do not persist | Local works, external is better |
| Cost at low traffic | Near zero | You pay for idle |
| Cost at steady traffic | Higher per request | Lower |
Three traps catch people specifically with agents.
Timeouts must nest. If the platform kills the function at 30 seconds and your agent's deadline is 45, the platform wins and you get a truncated request with no log line explaining why. Set the agent's cap below the platform's — 25 against 30 — so your code owns the failure and can log it.
Local state does not exist. An SQLite cache and an in-process token bucket both live on one instance. With ten concurrent instances you get ten independent caches (hit rate collapses) and ten independent rate limiters (a user can send ten times your intended rate). Move both to Redis before you scale out.
Size limits bite late. The full agent stack with HTML parsing and HTTP libraries can approach or exceed a 250 MB unzipped limit once transitive dependencies land. Check du -sh on your installed packages on day one, not in week six.
1# api/research.py - serverless entry point2from http.server import BaseHTTPRequestHandler3import json45_agent = build_agent()6_cache = AnswerCache(path="/tmp/cache.db") # per-instance; use Redis in production78class handler(BaseHTTPRequestHandler):9 def do_POST(self):10 body = json.loads(self.rfile.read(11 int(self.headers.get("Content-Length", 0))) or b"{}")12 q = (body.get("question") or "").strip()13 if not q or len(q) > 500:14 self._send(400, {"error": "question must be 1-500 characters"})15 return16 self._send(200, answer(q, body.get("user_id", "anon"), _cache, _agent))1718 def _send(self, code, payload):19 data = json.dumps(payload).encode()20 self.send_response(code)21 self.send_header("Content-Type", "application/json")22 self.send_header("Content-Length", str(len(data)))23 self.end_headers()24 self.wfile.write(data)Note the 500-character input cap. Without it, someone pastes a 40,000-character document into question and every one of your six agent steps carries it. That is one request costing several dollars, and it is trivially avoidable.
Knowing whether it actually works
The agent will look impressive on the questions you thought of while building it. That tells you nothing, because you unconsciously chose questions it handles. Build a fixed set of fifteen questions covering the cases you did not design for: a question with no answer, a question whose answer changed last week, a question where two sources disagree, a question needing arithmetic on two retrieved figures, and a question whose obvious search terms return junk.
Record four things per question: which tools ran, whether every figure in the answer appears in some tool result, whether the cited URLs appeared in tool results, and the latency. The second and third are the ones that matter — they detect hallucination mechanically, without a human reading the answer, and they are checkable with a substring test:
1def unsupported_numbers(answer_text: str, steps) -> list[str]:2 import re3 observations = " ".join(str(s[1]) for s in steps)4 figures = re.findall(r"\d[\d,]*\.?\d*", answer_text)5 return [f for f in figures if f.replace(",", "") not in6 observations.replace(",", "")]Any figure this returns is a figure the agent produced without evidence. On a well-built agent it returns an empty list almost always; when it does not, you have caught the exact failure that makes research agents dangerous — a confident, sourced-looking answer with an invented number in it. Run this on every response in production, log the failures, and you will know about the problem before a user does.