Course Content
AI Agent Frameworks
4 sections · 15 lessons
Exercise 1: Build a LangChain Web Query Agent
You are going to build a small agent that can do two things a language model cannot do alone: look things up on the live web, and perform arithmetic that is actually correct. Then you will make it remember the conversation, so that "and what about Ford?" works as a follow-up question.
That combination is deliberately chosen. It is the smallest system that exhibits every failure mode of tool-using agents. By the time it works reliably you will have hit — and fixed — the wrong-tool problem, the runaway-loop problem, the silent-hallucinated-number problem, and the amnesia problem. Those four account for most of what goes wrong in real deployments.
Here is the target behaviour. Type this, and get this:
You: What was Tesla's revenue in Q3 2024?Bot: Tesla reported $25.18 billion in revenue for Q3 2024. [search_web called once]You: What's 15% of that?Bot: 15% of $25.18 billion is $3.777 billion. [calculate called once with "0.15 * 25180000000"]You: And how does that compare to what I asked about first?Bot: You first asked about Tesla's Q3 2024 revenue of $25.18 billion... [no tools called - answered from memory]Three turns, three different behaviours: a search, a calculation that references the previous answer, and a memory lookup with no tool at all. Getting all three right is the exercise.
Setup
python -m venv .venv && source .venv/bin/activatepip install langchain langchain-anthropic langgraph ddgs python-dotenvCreate a .env file next to your script:
ANTHROPIC_API_KEY=sk-ant-...Never hard-code the key in the script. This is not a style point — API keys committed to git are scraped from public repositories within minutes, and the resulting bill is yours.
Starter skeleton
1import os2from dotenv import load_dotenv3from langchain_core.tools import tool4from langchain_anthropic import ChatAnthropic5from langchain.agents import create_agent67load_dotenv()89@tool10def search_web(query: str) -> str:11 """TODO: write a description that tells the model exactly when to use this."""12 raise NotImplementedError1314@tool15def calculate(expression: str) -> str:16 """TODO: write a description that tells the model exactly when to use this."""17 raise NotImplementedError1819def build_agent():20 raise NotImplementedError2122if __name__ == "__main__":23 agent = build_agent()24 while True:25 q = input("You: ").strip()26 if q.lower() in {"quit", "exit"}:27 break28 out = agent.invoke({"messages": [{"role": "user", "content": q}]})29 print("Bot:", out["messages"][-1].content)Requirement 1 — the search tool
Use DuckDuckGo through the ddgs package, which needs no API key.
1from ddgs import DDGS23@tool4def search_web(query: str) -> str:5 """Search the live web for current, factual information: news, prices,6 company financials, recent events, and people's current roles.78 Use this whenever the answer depends on facts that may have changed9 recently or that you are not certain of. Do NOT use it for arithmetic.1011 The query should be 2-8 keywords, no punctuation and no question marks.12 Good: 'Tesla Q3 2024 revenue'. Bad: 'What was Tesla revenue?'1314 Returns up to 5 numbered snippets, or a message saying nothing was found.15 """16 try:17 with DDGS() as ddgs:18 hits = list(ddgs.text(query, max_results=5))19 except Exception as e:20 return f"ERROR: search unavailable ({type(e).__name__}). Continue without it."2122 if not hits:23 return (f"No results for {query!r}. Try broader keywords once. "24 f"If it fails again, tell the user the information is unavailable.")2526 return "\n".join(27 f"{i}. {h['title']}\n {h['body'][:250]}\n {h['href']}"28 for i, h in enumerate(hits, 1)29 )Three deliberate design decisions in that docstring, each fixing a specific failure:
| Line | Prevents |
|---|---|
| "Do NOT use it for arithmetic" | The agent searching for "15% of 25.18 billion" |
| "2-8 keywords, no question marks" | Passing the raw user question, which returns forum threads |
| "If it fails again, tell the user" | The endless-retry loop when a query genuinely has no results |
Also note that failures return strings rather than raising. A raised exception ends the whole request; a returned string becomes the agent's next observation, and it can adapt.
Requirement 2 — the calculator
The obvious implementation is eval(expression). Do not write it. eval on text that a model produced from user input means a user who types "compute the result of __import__('os').system('rm -rf /')" may get exactly what they asked for. Parse instead.
1import ast, operator23_OPS = {ast.Add: operator.add, ast.Sub: operator.sub,4 ast.Mult: operator.mul, ast.Div: operator.truediv,5 ast.Pow: operator.pow, ast.USub: operator.neg,6 ast.Mod: operator.mod}78def _eval(node):9 if isinstance(node, ast.Constant) and isinstance(node.value, (int, float)):10 return node.value11 if isinstance(node, ast.BinOp) and type(node.op) in _OPS:12 return _OPS[type(node.op)](_eval(node.left), _eval(node.right))13 if isinstance(node, ast.UnaryOp) and type(node.op) in _OPS:14 return _OPS[type(node.op)](_eval(node.operand))15 raise ValueError("only numbers and + - * / % ** are allowed")1617@tool18def calculate(expression: str) -> str:19 """Evaluate an arithmetic expression exactly.2021 ALWAYS use this for every arithmetic step, including percentages,22 differences, ratios and growth rates - even ones that look easy.23 Never compute a number yourself.2425 Pass numbers in full, not abbreviated: write 25180000000, not 25.18B.26 Example: '0.15 * 25180000000'.2728 Returns the numeric result, or a message starting with ERROR:.29 """30 try:31 value = _eval(ast.parse(expression, mode="eval").body)32 except ZeroDivisionError:33 return "ERROR: division by zero."34 except Exception as e:35 return (f"ERROR: {e}. Send only digits, decimal points and "36 f"the operators + - * / % ** with brackets.")37 return f"{value:,.4f}".rstrip("0").rstrip(".")"Write 25180000000, not 25.18B" is the line that matters most. Without it, models routinely pass "0.15 * 25.18B", the parser rejects it, and the agent burns two extra steps recovering. With it, the first call is well-formed about nine times out of ten.
Every constraint your tool enforces must be stated in the tool description. A rule the model cannot see is a rule it will break, and each break costs you a full model call.
Requirement 3 — the agent with memory
1from langgraph.checkpoint.memory import InMemorySaver2from langchain.agents.middleware import ModelCallLimitMiddleware34SYSTEM = """You are a research assistant with two tools.56Rules you must follow exactly:71. Never state a number you did not receive from a tool result.82. Use `calculate` for every arithmetic step, however trivial.93. Use `search_web` only for facts that may have changed or that you are10 unsure of. Never search for something you have already found this session.114. If a search returns nothing useful twice, say the information is12 unavailable and stop searching.135. Answer in at most three sentences and name the source you used."""1415def build_agent():16 llm = ChatAnthropic(model=os.getenv("AGENT_MODEL", "claude-sonnet-5"))17 return create_agent(18 model=llm,19 tools=[search_web, calculate],20 system_prompt=SYSTEM,21 middleware=[ModelCallLimitMiddleware(run_limit=5, exit_behavior="end")],22 checkpointer=InMemorySaver(), # remembers earlier turns23 )2425def this_turn(messages):26 """The messages produced since the latest user message."""27 last_user = max(i for i, m in enumerate(messages) if m.type == "human")28 return messages[last_user + 1:]Two kinds of memory live in the same messages list, and confusing them is the classic mistake:
| Memory | Contains | Lifetime | If it is missing |
|---|---|---|---|
Tool calls and results in messages | What the agent did within this request | One request | Agent repeats the same tool call forever |
Checkpointer plus thread_id | Earlier user and assistant turns | The thread | "What about Ford?" fails — no antecedent |
Update the main loop to pass a thread ID, or the checkpointer has nowhere to store anything:
cfg = {"configurable": {"thread_id": "cli-session"}}out = agent.invoke({"messages": [{"role": "user", "content": q}]}, cfg)print("Bot:", out["messages"][-1].content)With a checkpointer, out["messages"] holds the whole thread, not just this turn. That is why this_turn exists: it cuts the list at the latest user message, so you can see which tools this question used.
Requirement 4 — test it against a fixed script
Do not test by improvising questions. Use the same six every time, so that when you change a docstring you can tell whether behaviour improved or regressed.
| # | Question | Correct tool sequence | Tests |
|---|---|---|---|
| 1 | What was Tesla's revenue in Q3 2024? | search_web | Search is chosen for a factual lookup |
| 2 | What's 15% of that? | calculate | Reference resolution plus tool switch |
| 3 | What is 847 * 293? | calculate | No search for pure arithmetic |
| 4 | Who is the current CEO of OpenAI? | search_web | Recency-sensitive fact |
| 5 | What is the capital of France? | none | Restraint — no tool needed |
| 6 | What did I ask about first? | none | Memory works |
Question 3 has a checkable answer: 847×293=248,171. If your agent returns anything else, it computed the number itself instead of calling the tool — which is exactly the failure rule 2 in the system prompt is meant to prevent. Question 5 is the opposite test: an agent that searches the web for the capital of France is over-eager, and that costs you a second of latency and a fraction of a cent on every trivial question.
Score the run on two axes separately, because they fail independently:
1SCRIPT = [2 ("What was Tesla's revenue in Q3 2024?", {"search_web"}),3 ("What's 15% of that?", {"calculate"}),4 ("What is 847 * 293?", {"calculate"}),5 ("Who is the current CEO of OpenAI?", {"search_web"}),6 ("What is the capital of France?", set()),7 ("What did I ask about first?", set()),8]910cfg = {"configurable": {"thread_id": "eval-run"}}11for q, expected in SCRIPT:12 out = agent.invoke({"messages": [{"role": "user", "content": q}]}, cfg)13 used = {m.name for m in this_turn(out["messages"]) if m.type == "tool"}14 verdict = "OK " if used == expected else "MISS"15 print(f"{verdict} {q}\n used={used or '{}'} expected={expected or '{}'}")When it does not work
| Symptom | Cause | Fix |
|---|---|---|
| Answers from the model's own knowledge instead of searching | Description does not say when to search | Add "Use this whenever the answer depends on facts that may have changed" |
| Searches for arithmetic | calculate's description is weaker than search_web's | Add "ALWAYS use this for every arithmetic step" and "Do NOT search for arithmetic" |
| Same search repeated 5+ times | Empty results returned as "" | Return a sentence granting permission to give up |
ERROR: only numbers and ... allowed repeatedly | Model sending 25.18B or 15% | Put "write 25180000000, not 25.18B" in the description |
| "What about Ford?" gets "Ford what?" | No checkpointer, or no thread_id in config | Add both; confirm with agent.get_state(cfg).values["messages"] |
| Memory works, then resets | Different thread_id per call, or process restarted | Fixed ID; persistent checkpointer for anything real |
| Hangs for a minute then errors | No deadline around the call | asyncio.wait_for(agent.ainvoke(...), timeout=40) |
When an agent picks the wrong tool, ninety per cent of the time the fix is in a docstring, not in Python. Change the description first and re-run the script before you change anything else.
Extensions worth doing
- Add a caching layer. Keep a dictionary from query string to result inside
search_web. On the six-question script it saves nothing; on a real session with repeated topics it removes a third of the network calls. - Add a
read_urltool that fetches one of the URLs the search returned and extracts its text. This makes the agent's reasoning two-stage: find candidates, then read one. Watch how much better the answers get, and how much slower. - Replace
InMemorySaverwithSqliteSaver(fromlanggraph-checkpoint-sqlite) so the conversation survives a restart. The change is one constructor. - Log every run as JSON — question, tools used, latency, final answer — and append to a file. After fifty runs you will have a dataset that tells you which question types your agent gets wrong.
What you should take away from getting this working
The code in this exercise is roughly 120 lines and almost none of it is interesting. The interesting part is the text: two docstrings and a five-rule system prompt. If you change nothing but those and re-run the six-question script, you will see the tool-selection score move — sometimes from 3/6 to 6/6 on the strength of one added sentence. That is the practical lesson, and it transfers to every agent you will ever build: the prompt surface is the program, and it deserves the same discipline as code. Version it, review it, and test it.
The second takeaway is about limits. A five-call model limit and a forty-second deadline look like defensive clutter until the day a search API starts returning empty results and your agent tries the same query forty times. With those two limits the worst case is five model calls and forty seconds. Without them the worst case is bounded only by your API quota.
The third is about verification. Get into the habit of looking at the ToolMessages in each result to see which tools actually ran. An agent that produces the right answer for the wrong reason — guessing a number that happens to be close — will pass a casual eyeball test and fail the moment the question changes slightly. The tool trace is the only place that distinction is visible.