Course Content
AI Agent Frameworks
4 sections · 15 lessons
LangChain Agents and Tool Execution
Here is a bug report that took a team two days to understand.
Their agent had eight tools. A user asked: "What's 15% of our Q3 revenue?" The agent called search_web, found nothing useful, called search_web again, then answered "15% of Q3 revenue is approximately 2.1 million dollars" — a number that appeared nowhere in any of its tool results. It had made it up.
The agent had a perfectly good calculate tool. It never touched it. The reason was one line of code:
1@tool2def calculate(expression: str) -> str:3 """Calculator."""4 return str(eval(expression))That docstring — the single word "Calculator." — is the entire description the model receives. It has no idea when to use it, what an expression looks like, or that it should prefer this over guessing. Meanwhile search_web's description was three sentences long and said "use this to find information." Faced with a vague tool and a well-described one, the model picked the well-described one. It behaved exactly as it was told.
This is the thing to internalise about LangChain agents before any API detail: the tool descriptions are the program. The Python code around them is scaffolding. The behaviour lives in the text.
What LangChain actually is
LangChain is a library for composing model calls with everything around them — prompts, tools, retrieval, memory, output parsing. Its agent module is one part of a larger toolkit, and it is the part that supplies the loop: give a model a set of tools, let it choose one, run it, feed the result back, repeat until the model produces a final answer instead of a tool call.
Since LangChain 1.0 (October 2025) that loop is built with one function, create_agent, which returns a small LangGraph graph under the hood. The older AgentExecutor and create_tool_calling_agent API that many tutorials still show now lives in a separate langchain-classic package, kept for existing code. This lesson uses the current API.
What you get for adopting it, concretely:
| You would otherwise write | LangChain gives you |
|---|---|
| Function signature → JSON schema conversion | @tool decorator reading type hints and docstrings |
| Provider-specific tool-call formats | One interface across Anthropic, OpenAI, Google, local models |
| Argument parsing and dispatch | The create_agent loop, with argument validation |
| Step limits, retries, error policy | Middleware: ModelCallLimitMiddleware, ToolCallLimitMiddleware, ToolRetryMiddleware |
| Conversation state plumbing | A LangGraph checkpointer plus a thread_id |
| Structured traces | Callback system, plus hosted tracing |
The trade you make: a create_agent run is three or four layers of abstraction — graph, nodes, middleware hooks — between your code and the HTTP request. When something goes wrong, you will spend time reading library source. That is the real cost, and it is worth paying only because the alternative is re-deriving the same twelve behaviours yourself.
Building an agent, properly
Installation
pip install langchain langchain-anthropic langgraphDefining tools that the model can actually use
A tool is a Python function plus a contract. The contract has four parts, and skipping any of them degrades behaviour measurably.
1from langchain_core.tools import tool2from pydantic import BaseModel, Field3import ast, operator45class CalcArgs(BaseModel):6 expression: str = Field(7 description="A pure arithmetic expression using only numbers and "8 "+ - * / ( ) and **. Example: '0.15 * 4820000'. "9 "No variables, no function names, no units."10 )1112@tool("calculate", args_schema=CalcArgs)13def calculate(expression: str) -> str:14 """Evaluate an arithmetic expression exactly.1516 ALWAYS use this for any arithmetic, including percentages, ratios,17 growth rates and currency conversion, instead of computing mentally.18 Mental arithmetic is unreliable; this is not.19 Returns the numeric result as a string, or an error message20 starting with 'ERROR:' if the expression is malformed.21 """22 allowed = {ast.Add: operator.add, ast.Sub: operator.sub,23 ast.Mult: operator.mul, ast.Div: operator.truediv,24 ast.Pow: operator.pow, ast.USub: operator.neg}2526 def ev(node):27 if isinstance(node, ast.Constant) and isinstance(node.value, (int, float)):28 return node.value29 if isinstance(node, ast.BinOp):30 return allowed[type(node.op)](ev(node.left), ev(node.right))31 if isinstance(node, ast.UnaryOp):32 return allowed[type(node.op)](ev(node.operand))33 raise ValueError("unsupported expression")3435 try:36 return str(ev(ast.parse(expression, mode="eval").body))37 except Exception as e:38 return (f"ERROR: could not evaluate {expression!r} ({e}). "39 f"Send a pure arithmetic expression with no words.")Four things are doing work here, and none of them are optional:
- The name is a verb phrase the model reads as an action.
- The docstring says when to use it and when not to. "ALWAYS use this for any arithmetic" is not decoration — it is the instruction that stops the model computing 15% in its head and getting it wrong.
- The argument schema shows an example. Models are dramatically better at producing well-formed arguments when the description contains one.
- The error return is a string, not an exception, and it tells the model how to fix the call. An exception ends the run; this string becomes the next observation and the model retries correctly.
Note also what the implementation does not use: eval(). A tool wired to eval() and exposed to user text is remote code execution with extra steps. The AST walker accepts arithmetic and nothing else.
Assembling the agent
1import os2from langchain_anthropic import ChatAnthropic3from langchain.agents import create_agent4from langchain.agents.middleware import (5 ModelCallLimitMiddleware, ToolCallLimitMiddleware)67llm = ChatAnthropic(model=os.getenv("AGENT_MODEL", "claude-sonnet-5"))89SYSTEM = ("You are a research assistant. Rules:\n"10 "1. Never state a number you did not obtain from a tool.\n"11 "2. Use `calculate` for every arithmetic step, however simple.\n"12 "3. If a search returns nothing useful twice, say so and stop.\n"13 "4. Cite the tool result you took each fact from.")1415agent = create_agent(16 model=llm,17 tools=[calculate, search_web],18 system_prompt=SYSTEM,19 middleware=[20 ModelCallLimitMiddleware(run_limit=6, exit_behavior="end"), # step cap21 ToolCallLimitMiddleware(tool_name="search_web", run_limit=3),22 ],23)2425result = agent.invoke(26 {"messages": [{"role": "user", "content": "What is 15% of 4,820,000?"}]})27print(result["messages"][-1].content)2829def tool_steps(messages):30 """(tool name, result text) for every tool call in a run, in order."""31 return [(m.name, m.content) for m in messages if m.type == "tool"]The model name comes from configuration with a current default, so changing models is an environment change, not a code change. There is no temperature=0 here on purpose: some current models reject or ignore sampling settings, so do not treat it as a universal switch.
The agent's state is one list, messages. Each step appends the model's reply, with its tool calls, and one ToolMessage per tool result, and the whole list is sent back to the model on the next step. That running list is what older LangChain code called the scratchpad: it is how the agent knows what it just did. tool_steps reads the tool results back out of it, which is how you will check which tools a run actually used.
The loop, step by step
Trace one request in full so there is no mystery about what the agent loop does.
User: "What was Tesla's Q3 2024 revenue, and what is 15% of it?"Step 1 send: system + user message model returns tool_call: search_web(query="Tesla Q3 2024 revenue") tool node runs it -> "Tesla reported $25.18B revenue in Q3 2024." messages now hold one (call, result) pairStep 2 send: system + user + messages(1 pair) model returns tool_call: calculate(expression="0.15 * 25180000000") tool node runs it -> "3777000000.0" messages now hold two pairsStep 3 send: system + user + messages(2 pairs) model returns NO tool call, just text the loop stops and the text is the final answerThree model calls, two tool calls. The rule the loop follows is simply: reply contains a tool call → run it, append, loop. Reply contains no tool call → that is the answer, stop. Everything else — iteration counting, timeouts, error conversion — wraps that rule.
The agent does not "decide to use a tool" in any deep sense. It emits a structured tool call because the schema was in its context and the description matched the situation. Better descriptions, better decisions.
Older ReAct-style prompting versus native tool calling
The original ReAct pattern (reason, act, observe) made the model produce text in a fixed shape which a parser then scraped:
Thought: I need Tesla's revenue.Action: search_webAction Input: Tesla Q3 2024 revenueObservation: Tesla reported $25.18B revenue in Q3 2024.Thought: Now I compute 15%.Action: calculateAction Input: 0.15 * 25180000000This text-parsing agent still exists in the legacy langchain-classic package as create_react_agent, and it is only worth reaching for with a model that has no native tool-calling API. But it is strictly worse when native tool calling is available, because the parse can fail: a model that writes Action Input: 0.15 × 25.18B breaks the regex, and now you are debugging string parsing rather than logic. Native tool calling returns structured JSON that cannot be malformed in that way.
| Text ReAct | Native tool calling | |
|---|---|---|
| Model requirement | Any instruction-following model | Model with tool-use API |
| Failure mode | Unparseable output | Invalid arguments (schema-caught) |
| Parallel tool calls | No | Yes, several per step |
| Token overhead | Higher — format examples in prompt | Lower — schemas sent natively |
| Use when | Local or older models | Everything else |
Adding memory across turns
By default the messages list starts empty on every request. To let the agent remember earlier turns, give it a checkpointer, which saves the state after every step, and pass a thread_id that names the conversation.
1from langgraph.checkpoint.memory import InMemorySaver23chat_agent = create_agent(model=llm, tools=[calculate, search_web],4 system_prompt=SYSTEM, checkpointer=InMemorySaver())56cfg = {"configurable": {"thread_id": "user-42"}}7chat_agent.invoke({"messages": [{"role": "user",8 "content": "What was Tesla's Q3 2024 revenue?"}]}, cfg)9chat_agent.invoke({"messages": [{"role": "user",10 "content": "And 15% of that?"}]}, cfg) # resolves "that"Two warnings. First, InMemorySaver is a dictionary in your process — it dies on restart and does not exist for the other pod behind your load balancer. For anything real, use a database-backed checkpointer such as PostgresSaver. Second, unbounded history is a cost bomb. A 40-turn conversation where each turn is 600 tokens sends 24,000 tokens of history on turn 41 alone. Summarise or trim it with middleware:
1from langchain.agents.middleware import SummarizationMiddleware23chat_agent = create_agent(4 model=llm, tools=[calculate, search_web], system_prompt=SYSTEM,5 checkpointer=InMemorySaver(),6 middleware=[SummarizationMiddleware(7 model="anthropic:claude-haiku-4-5", # a cheap model writes the summary8 trigger=("tokens", 3000), # summarise once history passes this9 keep=("messages", 10))], # keep the latest 10 messages verbatim10)What happens when a tool fails
There are three distinct failure classes and they need three different responses. Conflating them is the most common source of agents that either crash constantly or loop forever.
| Failure | Example | Right response | Why |
|---|---|---|---|
| Model's fault | calculate("fifteen percent") | Return an ERROR: string explaining the correct format | The model can fix this on the next step |
| Transient | API returns 503 or times out | Retry inside the tool with backoff, then return a string | Retrying is cheaper than a whole extra model call |
| Fatal | Invalid API key, quota exhausted | Raise — let the run die | No amount of model cleverness fixes a missing key |
1import time, requests23@tool4def search_web(query: str) -> str:5 """Search the live web for current facts: news, prices, recent events,6 people's current roles. Not for arithmetic and not for general knowledge7 the model already has. Returns up to 5 result snippets."""8 for attempt in range(3):9 try:10 r = requests.get(API, params={"q": query}, timeout=8)11 if r.status_code == 429:12 time.sleep(2 ** attempt) # 1s, 2s, 4s13 continue14 r.raise_for_status()15 hits = r.json().get("results", [])[:5]16 if not hits:17 return (f"No results for {query!r}. Try broader terms, "18 f"or state that the information is unavailable.")19 return "\n".join(f"- {h['title']}: {h['snippet']}" for h in hits)20 except requests.Timeout:21 if attempt == 2:22 return "ERROR: search timed out three times. Proceed without it."23 except requests.HTTPError as e:24 if e.response.status_code in (401, 403):25 raise # fatal: bad credentials26 return f"ERROR: search failed ({e.response.status_code})."27 return "ERROR: search unavailable."Notice the empty-results branch. Returning "" for no results is what caused the forty-one-search loop in the opening story: an empty string looks to the model like "something went slightly wrong, try again". A sentence that explicitly says "or state that the information is unavailable" gives it a legitimate exit.
Decide, for every tool, whether its failure is the model's problem or yours. Return a string when the model can fix it; raise only when no amount of cleverness can.
Debugging: seeing what the model saw
Streaming the steps
1question = {"messages": [{"role": "user",2 "content": "Compare Tesla and Ford Q3 revenue"}]}3for chunk in agent.stream(question, stream_mode="updates"):4 for node, update in chunk.items():5 for m in (update or {}).get("messages", []):6 if m.type == "ai" and m.tool_calls:7 for c in m.tool_calls:8 print(f"CALL {c['name']}({c['args']})")9 elif m.type == "tool":10 print(f"RESULT {m.content[:160]}")11 else:12 print(f"ANSWER {m.content}")Callbacks for structured traces
1from langchain_core.callbacks import BaseCallbackHandler2import time, json, uuid34class TraceHandler(BaseCallbackHandler):5 def __init__(self):6 self.run_id = str(uuid.uuid4()); self.t0 = {}; self.calls = 078 def on_tool_start(self, serialized, input_str, **kw):9 name = serialized.get("name", "?")10 self.t0[name] = time.time(); self.calls += 11112 def on_tool_end(self, output, **kw):13 name = kw.get("name", "?")14 print(json.dumps({"run": self.run_id, "tool": name,15 "ms": round((time.time() - self.t0.get(name, 0)) * 1000),16 "chars": len(str(output))}))1718 def on_llm_end(self, response, **kw):19 u = response.generations[0][0].message.usage_metadata20 print(json.dumps({"run": self.run_id, "usage": u}))2122agent.invoke({"messages": [{"role": "user", "content": "..."}]},23 config={"callbacks": [TraceHandler()]})Setting the environment variables LANGSMITH_TRACING=true and LANGSMITH_API_KEY=... sends the same information to a hosted trace viewer, which is worth it once traces get long enough that reading JSON in a terminal stops working.
The four ways LangChain agents go wrong
Vague tool descriptions
Covered above, and it remains the single largest lever. A useful discipline: write the description, then read it as if you were a contractor who has never seen the codebase. If you could not tell from that text alone when to call it and what to pass, the model cannot either.
Exceptions instead of error strings
A tool that raises on a 404 turns a recoverable situation into a dead request. Wrap the body, return a string starting with ERROR:, and reserve raising for genuinely fatal conditions.
Too many tools
Tool selection accuracy falls as the tool count rises, and the schemas cost tokens on every single step. Suppose each tool schema is about 180 tokens. With 5 tools that is 900 tokens per step; over a 6-step run, 5,400 tokens. With 25 tools it is 4,500 per step and 27,000 per run — an extra 0.065 dollars per request at an example rate of 3 dollars per million tokens, before the model has done anything useful. Worse, the model now has 25 plausible options and picks wrong more often.
The fix is routing: a cheap first classification picks a subset of tools, and the agent is constructed with only those.
1TOOLSETS = {2 "finance": [get_stock_price, get_financials, calculate],3 "research": [search_web, read_url, summarise],4 "internal": [query_db, lookup_order, calculate],5}67def route(question: str) -> str:8 label = classifier_llm.invoke(9 f"Reply with exactly one word - finance, research or internal:\n{question}"10 ).content.strip().lower()11 return label if label in TOOLSETS else "research"1213tools = TOOLSETS[route(user_question)] # 3 tools in context, not 25Unbounded tool use
Always set a model-call limit and a wall-clock deadline. Without them the only stop is LangGraph's recursion limit, which by default is in the thousands of steps. A cap of 6 model calls with a 45-second deadline bounds a runaway request to roughly 6 model calls instead of 41, and turns a failure costing tens of dollars into one costing well under a dollar. create_agent has no deadline argument, so put one around the call — await asyncio.wait_for(agent.ainvoke(...), timeout=45). Then handle the cap explicitly rather than shipping the terse "Model call limits exceeded" message to users:
1out = agent.invoke({"messages": [{"role": "user", "content": q}]})2model_turns = sum(m.type == "ai" for m in out["messages"])3answer = out["messages"][-1].content4if model_turns >= 6:5 log.warning("hit model-call cap", extra={"q": q, "tools":6 [name for name, _ in tool_steps(out["messages"])]})7 found = "\n".join(f"- {r[:200]}" for _, r in tool_steps(out["messages"]))8 answer = ("I could not finish this within my step budget. "9 "Here is what I found so far:\n" + found)Making this survive contact with users
When you put a LangChain agent in front of real traffic, three habits separate the systems that hold up from the ones that generate support tickets.
Treat tool descriptions as versioned artefacts. They are prompt engineering, they change behaviour, and a one-word edit can change your tool-selection rate by tens of percentage points. Keep them in code review. When an agent misbehaves, diff the description before you touch anything else.
Build a fixed evaluation set of twenty questions with known correct tool sequences. For each, record which tools should be called. Run it after every change to a prompt, a description or a model version. Score two things separately: did it call the right tools, and was the final answer right. They fail independently, and knowing which one broke tells you where to look.
Assume every number in a final answer is a hallucination until a tool result proves otherwise. The opening bug — a confident "2.1 million dollars" from nowhere — is the characteristic failure of tool-using agents, and it is invisible unless you check. Keep the tool results — the ToolMessages in the returned messages — and for numeric answers, verify programmatically that each figure appears in some tool result. If it does not, do not show it to the user.