AI Agent Frameworks

Exercise 1: Build a LangChain Web Query Agent


You are going to build a small agent that can do two things a language model cannot do alone: look things up on the live web, and perform arithmetic that is actually correct. Then you will make it remember the conversation, so that "and what about Ford?" works as a follow-up question.

That combination is deliberately chosen. It is the smallest system that exhibits every failure mode of tool-using agents. By the time it works reliably you will have hit — and fixed — the wrong-tool problem, the runaway-loop problem, the silent-hallucinated-number problem, and the amnesia problem. Those four account for most of what goes wrong in real deployments.

Here is the target behaviour. Type this, and get this:

Text
You: What was Tesla's revenue in Q3 2024?Bot: Tesla reported $25.18 billion in revenue for Q3 2024.     [search_web called once]You: What's 15% of that?Bot: 15% of $25.18 billion is $3.777 billion.     [calculate called once with "0.15 * 25180000000"]You: And how does that compare to what I asked about first?Bot: You first asked about Tesla's Q3 2024 revenue of $25.18 billion...     [no tools called - answered from memory]

Three turns, three different behaviours: a search, a calculation that references the previous answer, and a memory lookup with no tool at all. Getting all three right is the exercise.

Building up to "and what about Ford?"Skeleton with amodel and a loopAdd a liveweb search toolAdd a calculatorfor real arithmeticAttach memoryacross turnsReplay afixedtest scriptThe follow-up only resolves if the earlier turns are still in the message list.
Search and arithmetic fix what the model cannot do; memory is what makes a pronoun resolve.

Setup

Bash
python -m venv .venv && source .venv/bin/activatepip install langchain langchain-anthropic langgraph ddgs python-dotenv

Create a .env file next to your script:

Bash
ANTHROPIC_API_KEY=sk-ant-...

Never hard-code the key in the script. This is not a style point — API keys committed to git are scraped from public repositories within minutes, and the resulting bill is yours.

Starter skeleton

Python
import osfrom dotenv import load_dotenvfrom langchain_core.tools import toolfrom langchain_anthropic import ChatAnthropicfrom langchain.agents import create_agentload_dotenv()@tooldef search_web(query: str) -> str:    """TODO: write a description that tells the model exactly when to use this."""    raise NotImplementedError@tooldef calculate(expression: str) -> str:    """TODO: write a description that tells the model exactly when to use this."""    raise NotImplementedErrordef build_agent():    raise NotImplementedErrorif __name__ == "__main__":    agent = build_agent()    while True:        q = input("You: ").strip()        if q.lower() in {"quit", "exit"}:            break        out = agent.invoke({"messages": [{"role": "user", "content": q}]})        print("Bot:", out["messages"][-1].content)

Requirement 1 — the search tool

Use DuckDuckGo through the ddgs package, which needs no API key.

Python
from ddgs import DDGS@tooldef search_web(query: str) -> str:    """Search the live web for current, factual information: news, prices,    company financials, recent events, and people's current roles.    Use this whenever the answer depends on facts that may have changed    recently or that you are not certain of. Do NOT use it for arithmetic.    The query should be 2-8 keywords, no punctuation and no question marks.    Good: 'Tesla Q3 2024 revenue'. Bad: 'What was Tesla revenue?'    Returns up to 5 numbered snippets, or a message saying nothing was found.    """    try:        with DDGS() as ddgs:            hits = list(ddgs.text(query, max_results=5))    except Exception as e:        return f"ERROR: search unavailable ({type(e).__name__}). Continue without it."    if not hits:        return (f"No results for {query!r}. Try broader keywords once. "                f"If it fails again, tell the user the information is unavailable.")    return "\n".join(        f"{i}. {h['title']}\n   {h['body'][:250]}\n   {h['href']}"        for i, h in enumerate(hits, 1)    )

Three deliberate design decisions in that docstring, each fixing a specific failure:

LinePrevents
"Do NOT use it for arithmetic"The agent searching for "15% of 25.18 billion"
"2-8 keywords, no question marks"Passing the raw user question, which returns forum threads
"If it fails again, tell the user"The endless-retry loop when a query genuinely has no results

Also note that failures return strings rather than raising. A raised exception ends the whole request; a returned string becomes the agent's next observation, and it can adapt.

Requirement 2 — the calculator

The obvious implementation is eval(expression). Do not write it. eval on text that a model produced from user input means a user who types "compute the result of __import__('os').system('rm -rf /')" may get exactly what they asked for. Parse instead.

Python
import ast, operator_OPS = {ast.Add: operator.add, ast.Sub: operator.sub,        ast.Mult: operator.mul, ast.Div: operator.truediv,        ast.Pow: operator.pow, ast.USub: operator.neg,        ast.Mod: operator.mod}def _eval(node):    if isinstance(node, ast.Constant) and isinstance(node.value, (int, float)):        return node.value    if isinstance(node, ast.BinOp) and type(node.op) in _OPS:        return _OPS[type(node.op)](_eval(node.left), _eval(node.right))    if isinstance(node, ast.UnaryOp) and type(node.op) in _OPS:        return _OPS[type(node.op)](_eval(node.operand))    raise ValueError("only numbers and + - * / % ** are allowed")@tooldef calculate(expression: str) -> str:    """Evaluate an arithmetic expression exactly.    ALWAYS use this for every arithmetic step, including percentages,    differences, ratios and growth rates - even ones that look easy.    Never compute a number yourself.    Pass numbers in full, not abbreviated: write 25180000000, not 25.18B.    Example: '0.15 * 25180000000'.    Returns the numeric result, or a message starting with ERROR:.    """    try:        value = _eval(ast.parse(expression, mode="eval").body)    except ZeroDivisionError:        return "ERROR: division by zero."    except Exception as e:        return (f"ERROR: {e}. Send only digits, decimal points and "                f"the operators + - * / % ** with brackets.")    return f"{value:,.4f}".rstrip("0").rstrip(".")

"Write 25180000000, not 25.18B" is the line that matters most. Without it, models routinely pass "0.15 * 25.18B", the parser rejects it, and the agent burns two extra steps recovering. With it, the first call is well-formed about nine times out of ten.

Every constraint your tool enforces must be stated in the tool description. A rule the model cannot see is a rule it will break, and each break costs you a full model call.

Requirement 3 — the agent with memory

Python
from langgraph.checkpoint.memory import InMemorySaverfrom langchain.agents.middleware import ModelCallLimitMiddlewareSYSTEM = """You are a research assistant with two tools.Rules you must follow exactly:1. Never state a number you did not receive from a tool result.2. Use `calculate` for every arithmetic step, however trivial.3. Use `search_web` only for facts that may have changed or that you are   unsure of. Never search for something you have already found this session.4. If a search returns nothing useful twice, say the information is   unavailable and stop searching.5. Answer in at most three sentences and name the source you used."""def build_agent():    llm = ChatAnthropic(model=os.getenv("AGENT_MODEL", "claude-sonnet-5"))    return create_agent(        model=llm,        tools=[search_web, calculate],        system_prompt=SYSTEM,        middleware=[ModelCallLimitMiddleware(run_limit=5, exit_behavior="end")],        checkpointer=InMemorySaver(),       # remembers earlier turns    )def this_turn(messages):    """The messages produced since the latest user message."""    last_user = max(i for i, m in enumerate(messages) if m.type == "human")    return messages[last_user + 1:]

Two kinds of memory live in the same messages list, and confusing them is the classic mistake:

MemoryContainsLifetimeIf it is missing
Tool calls and results in messagesWhat the agent did within this requestOne requestAgent repeats the same tool call forever
Checkpointer plus thread_idEarlier user and assistant turnsThe thread"What about Ford?" fails — no antecedent

Update the main loop to pass a thread ID, or the checkpointer has nowhere to store anything:

Python
cfg = {"configurable": {"thread_id": "cli-session"}}out = agent.invoke({"messages": [{"role": "user", "content": q}]}, cfg)print("Bot:", out["messages"][-1].content)

With a checkpointer, out["messages"] holds the whole thread, not just this turn. That is why this_turn exists: it cuts the list at the latest user message, so you can see which tools this question used.

Requirement 4 — test it against a fixed script

Do not test by improvising questions. Use the same six every time, so that when you change a docstring you can tell whether behaviour improved or regressed.

#QuestionCorrect tool sequenceTests
1What was Tesla's revenue in Q3 2024?search_webSearch is chosen for a factual lookup
2What's 15% of that?calculateReference resolution plus tool switch
3What is 847 * 293?calculateNo search for pure arithmetic
4Who is the current CEO of OpenAI?search_webRecency-sensitive fact
5What is the capital of France?noneRestraint — no tool needed
6What did I ask about first?noneMemory works

Question 3 has a checkable answer: 847×293=248,171847 \times 293 = 248{,}171. If your agent returns anything else, it computed the number itself instead of calling the tool — which is exactly the failure rule 2 in the system prompt is meant to prevent. Question 5 is the opposite test: an agent that searches the web for the capital of France is over-eager, and that costs you a second of latency and a fraction of a cent on every trivial question.

Score the run on two axes separately, because they fail independently:

Python
SCRIPT = [    ("What was Tesla's revenue in Q3 2024?", {"search_web"}),    ("What's 15% of that?",                  {"calculate"}),    ("What is 847 * 293?",                   {"calculate"}),    ("Who is the current CEO of OpenAI?",    {"search_web"}),    ("What is the capital of France?",       set()),    ("What did I ask about first?",          set()),]cfg = {"configurable": {"thread_id": "eval-run"}}for q, expected in SCRIPT:    out = agent.invoke({"messages": [{"role": "user", "content": q}]}, cfg)    used = {m.name for m in this_turn(out["messages"]) if m.type == "tool"}    verdict = "OK  " if used == expected else "MISS"    print(f"{verdict} {q}\n      used={used or '{}'} expected={expected or '{}'}")

When it does not work

SymptomCauseFix
Answers from the model's own knowledge instead of searchingDescription does not say when to searchAdd "Use this whenever the answer depends on facts that may have changed"
Searches for arithmeticcalculate's description is weaker than search_web'sAdd "ALWAYS use this for every arithmetic step" and "Do NOT search for arithmetic"
Same search repeated 5+ timesEmpty results returned as ""Return a sentence granting permission to give up
ERROR: only numbers and ... allowed repeatedlyModel sending 25.18B or 15%Put "write 25180000000, not 25.18B" in the description
"What about Ford?" gets "Ford what?"No checkpointer, or no thread_id in configAdd both; confirm with agent.get_state(cfg).values["messages"]
Memory works, then resetsDifferent thread_id per call, or process restartedFixed ID; persistent checkpointer for anything real
Hangs for a minute then errorsNo deadline around the callasyncio.wait_for(agent.ainvoke(...), timeout=40)

When an agent picks the wrong tool, ninety per cent of the time the fix is in a docstring, not in Python. Change the description first and re-run the script before you change anything else.

Extensions worth doing

  • Add a caching layer. Keep a dictionary from query string to result inside search_web. On the six-question script it saves nothing; on a real session with repeated topics it removes a third of the network calls.
  • Add a read_url tool that fetches one of the URLs the search returned and extracts its text. This makes the agent's reasoning two-stage: find candidates, then read one. Watch how much better the answers get, and how much slower.
  • Replace InMemorySaver with SqliteSaver (from langgraph-checkpoint-sqlite) so the conversation survives a restart. The change is one constructor.
  • Log every run as JSON — question, tools used, latency, final answer — and append to a file. After fifty runs you will have a dataset that tells you which question types your agent gets wrong.

What you should take away from getting this working

The code in this exercise is roughly 120 lines and almost none of it is interesting. The interesting part is the text: two docstrings and a five-rule system prompt. If you change nothing but those and re-run the six-question script, you will see the tool-selection score move — sometimes from 3/6 to 6/6 on the strength of one added sentence. That is the practical lesson, and it transfers to every agent you will ever build: the prompt surface is the program, and it deserves the same discipline as code. Version it, review it, and test it.

The second takeaway is about limits. A five-call model limit and a forty-second deadline look like defensive clutter until the day a search API starts returning empty results and your agent tries the same query forty times. With those two limits the worst case is five model calls and forty seconds. Without them the worst case is bounded only by your API quota.

The third is about verification. Get into the habit of looking at the ToolMessages in each result to see which tools actually ran. An agent that produces the right answer for the wrong reason — guessing a number that happens to be close — will pass a casual eyeball test and fail the moment the question changes slightly. The tool trace is the only place that distinction is visible.