AI Agent Frameworks

LangChain Agents and Tool Execution


Here is a bug report that took a team two days to understand.

Their agent had eight tools. A user asked: "What's 15% of our Q3 revenue?" The agent called search_web, found nothing useful, called search_web again, then answered "15% of Q3 revenue is approximately 2.1 million dollars" — a number that appeared nowhere in any of its tool results. It had made it up.

The agent had a perfectly good calculate tool. It never touched it. The reason was one line of code:

Python
@tooldef calculate(expression: str) -> str:    """Calculator."""    return str(eval(expression))

That docstring — the single word "Calculator." — is the entire description the model receives. It has no idea when to use it, what an expression looks like, or that it should prefer this over guessing. Meanwhile search_web's description was three sentences long and said "use this to find information." Faced with a vague tool and a well-described one, the model picked the well-described one. It behaved exactly as it was told.

This is the thing to internalise about LangChain agents before any API detail: the tool descriptions are the program. The Python code around them is scaffolding. The behaviour lives in the text.

One turn through a LangChain agentuser inputmodelnames a tooltool runsobservationappendedfinal answerdescriptiondecideserrorsreturn hereA tool that raises kills the run; a tool that returns its error lets the model try another route.
The tool description is the prompt the model actually reads — vague wording is a routing bug, not a docs bug.

What LangChain actually is

LangChain is a library for composing model calls with everything around them — prompts, tools, retrieval, memory, output parsing. Its agent module is one part of a larger toolkit, and it is the part that supplies the loop: give a model a set of tools, let it choose one, run it, feed the result back, repeat until the model produces a final answer instead of a tool call.

Since LangChain 1.0 (October 2025) that loop is built with one function, create_agent, which returns a small LangGraph graph under the hood. The older AgentExecutor and create_tool_calling_agent API that many tutorials still show now lives in a separate langchain-classic package, kept for existing code. This lesson uses the current API.

What you get for adopting it, concretely:

You would otherwise writeLangChain gives you
Function signature → JSON schema conversion@tool decorator reading type hints and docstrings
Provider-specific tool-call formatsOne interface across Anthropic, OpenAI, Google, local models
Argument parsing and dispatchThe create_agent loop, with argument validation
Step limits, retries, error policyMiddleware: ModelCallLimitMiddleware, ToolCallLimitMiddleware, ToolRetryMiddleware
Conversation state plumbingA LangGraph checkpointer plus a thread_id
Structured tracesCallback system, plus hosted tracing

The trade you make: a create_agent run is three or four layers of abstraction — graph, nodes, middleware hooks — between your code and the HTTP request. When something goes wrong, you will spend time reading library source. That is the real cost, and it is worth paying only because the alternative is re-deriving the same twelve behaviours yourself.

Building an agent, properly

Installation

Bash
pip install langchain langchain-anthropic langgraph

Defining tools that the model can actually use

A tool is a Python function plus a contract. The contract has four parts, and skipping any of them degrades behaviour measurably.

Python
from langchain_core.tools import toolfrom pydantic import BaseModel, Fieldimport ast, operatorclass CalcArgs(BaseModel):    expression: str = Field(        description="A pure arithmetic expression using only numbers and "                    "+ - * / ( ) and **. Example: '0.15 * 4820000'. "                    "No variables, no function names, no units."    )@tool("calculate", args_schema=CalcArgs)def calculate(expression: str) -> str:    """Evaluate an arithmetic expression exactly.    ALWAYS use this for any arithmetic, including percentages, ratios,    growth rates and currency conversion, instead of computing mentally.    Mental arithmetic is unreliable; this is not.    Returns the numeric result as a string, or an error message    starting with 'ERROR:' if the expression is malformed.    """    allowed = {ast.Add: operator.add, ast.Sub: operator.sub,               ast.Mult: operator.mul, ast.Div: operator.truediv,               ast.Pow: operator.pow, ast.USub: operator.neg}    def ev(node):        if isinstance(node, ast.Constant) and isinstance(node.value, (int, float)):            return node.value        if isinstance(node, ast.BinOp):            return allowed[type(node.op)](ev(node.left), ev(node.right))        if isinstance(node, ast.UnaryOp):            return allowed[type(node.op)](ev(node.operand))        raise ValueError("unsupported expression")    try:        return str(ev(ast.parse(expression, mode="eval").body))    except Exception as e:        return (f"ERROR: could not evaluate {expression!r} ({e}). "                f"Send a pure arithmetic expression with no words.")

Four things are doing work here, and none of them are optional:

  • The name is a verb phrase the model reads as an action.
  • The docstring says when to use it and when not to. "ALWAYS use this for any arithmetic" is not decoration — it is the instruction that stops the model computing 15% in its head and getting it wrong.
  • The argument schema shows an example. Models are dramatically better at producing well-formed arguments when the description contains one.
  • The error return is a string, not an exception, and it tells the model how to fix the call. An exception ends the run; this string becomes the next observation and the model retries correctly.

Note also what the implementation does not use: eval(). A tool wired to eval() and exposed to user text is remote code execution with extra steps. The AST walker accepts arithmetic and nothing else.

Assembling the agent

Python
import osfrom langchain_anthropic import ChatAnthropicfrom langchain.agents import create_agentfrom langchain.agents.middleware import (    ModelCallLimitMiddleware, ToolCallLimitMiddleware)llm = ChatAnthropic(model=os.getenv("AGENT_MODEL", "claude-sonnet-5"))SYSTEM = ("You are a research assistant. Rules:\n"          "1. Never state a number you did not obtain from a tool.\n"          "2. Use `calculate` for every arithmetic step, however simple.\n"          "3. If a search returns nothing useful twice, say so and stop.\n"          "4. Cite the tool result you took each fact from.")agent = create_agent(    model=llm,    tools=[calculate, search_web],    system_prompt=SYSTEM,    middleware=[        ModelCallLimitMiddleware(run_limit=6, exit_behavior="end"),  # step cap        ToolCallLimitMiddleware(tool_name="search_web", run_limit=3),    ],)result = agent.invoke(    {"messages": [{"role": "user", "content": "What is 15% of 4,820,000?"}]})print(result["messages"][-1].content)def tool_steps(messages):    """(tool name, result text) for every tool call in a run, in order."""    return [(m.name, m.content) for m in messages if m.type == "tool"]

The model name comes from configuration with a current default, so changing models is an environment change, not a code change. There is no temperature=0 here on purpose: some current models reject or ignore sampling settings, so do not treat it as a universal switch.

The agent's state is one list, messages. Each step appends the model's reply, with its tool calls, and one ToolMessage per tool result, and the whole list is sent back to the model on the next step. That running list is what older LangChain code called the scratchpad: it is how the agent knows what it just did. tool_steps reads the tool results back out of it, which is how you will check which tools a run actually used.

The loop, step by step

Trace one request in full so there is no mystery about what the agent loop does.

Text
User: "What was Tesla's Q3 2024 revenue, and what is 15% of it?"Step 1  send: system + user message        model returns tool_call: search_web(query="Tesla Q3 2024 revenue")        tool node runs it -> "Tesla reported $25.18B revenue in Q3 2024."        messages now hold one (call, result) pairStep 2  send: system + user + messages(1 pair)        model returns tool_call: calculate(expression="0.15 * 25180000000")        tool node runs it -> "3777000000.0"        messages now hold two pairsStep 3  send: system + user + messages(2 pairs)        model returns NO tool call, just text        the loop stops and the text is the final answer

Three model calls, two tool calls. The rule the loop follows is simply: reply contains a tool call → run it, append, loop. Reply contains no tool call → that is the answer, stop. Everything else — iteration counting, timeouts, error conversion — wraps that rule.

The agent does not "decide to use a tool" in any deep sense. It emits a structured tool call because the schema was in its context and the description matched the situation. Better descriptions, better decisions.

Older ReAct-style prompting versus native tool calling

The original ReAct pattern (reason, act, observe) made the model produce text in a fixed shape which a parser then scraped:

Text
Thought: I need Tesla's revenue.Action: search_webAction Input: Tesla Q3 2024 revenueObservation: Tesla reported $25.18B revenue in Q3 2024.Thought: Now I compute 15%.Action: calculateAction Input: 0.15 * 25180000000

This text-parsing agent still exists in the legacy langchain-classic package as create_react_agent, and it is only worth reaching for with a model that has no native tool-calling API. But it is strictly worse when native tool calling is available, because the parse can fail: a model that writes Action Input: 0.15 × 25.18B breaks the regex, and now you are debugging string parsing rather than logic. Native tool calling returns structured JSON that cannot be malformed in that way.

Text ReActNative tool calling
Model requirementAny instruction-following modelModel with tool-use API
Failure modeUnparseable outputInvalid arguments (schema-caught)
Parallel tool callsNoYes, several per step
Token overheadHigher — format examples in promptLower — schemas sent natively
Use whenLocal or older modelsEverything else

Adding memory across turns

By default the messages list starts empty on every request. To let the agent remember earlier turns, give it a checkpointer, which saves the state after every step, and pass a thread_id that names the conversation.

Python
from langgraph.checkpoint.memory import InMemorySaverchat_agent = create_agent(model=llm, tools=[calculate, search_web],                          system_prompt=SYSTEM, checkpointer=InMemorySaver())cfg = {"configurable": {"thread_id": "user-42"}}chat_agent.invoke({"messages": [{"role": "user",                   "content": "What was Tesla's Q3 2024 revenue?"}]}, cfg)chat_agent.invoke({"messages": [{"role": "user",                   "content": "And 15% of that?"}]}, cfg)   # resolves "that"

Two warnings. First, InMemorySaver is a dictionary in your process — it dies on restart and does not exist for the other pod behind your load balancer. For anything real, use a database-backed checkpointer such as PostgresSaver. Second, unbounded history is a cost bomb. A 40-turn conversation where each turn is 600 tokens sends 24,000 tokens of history on turn 41 alone. Summarise or trim it with middleware:

Python
from langchain.agents.middleware import SummarizationMiddlewarechat_agent = create_agent(    model=llm, tools=[calculate, search_web], system_prompt=SYSTEM,    checkpointer=InMemorySaver(),    middleware=[SummarizationMiddleware(        model="anthropic:claude-haiku-4-5",   # a cheap model writes the summary        trigger=("tokens", 3000),             # summarise once history passes this        keep=("messages", 10))],              # keep the latest 10 messages verbatim)

What happens when a tool fails

There are three distinct failure classes and they need three different responses. Conflating them is the most common source of agents that either crash constantly or loop forever.

FailureExampleRight responseWhy
Model's faultcalculate("fifteen percent")Return an ERROR: string explaining the correct formatThe model can fix this on the next step
TransientAPI returns 503 or times outRetry inside the tool with backoff, then return a stringRetrying is cheaper than a whole extra model call
FatalInvalid API key, quota exhaustedRaise — let the run dieNo amount of model cleverness fixes a missing key
Python
import time, requests@tooldef search_web(query: str) -> str:    """Search the live web for current facts: news, prices, recent events,    people's current roles. Not for arithmetic and not for general knowledge    the model already has. Returns up to 5 result snippets."""    for attempt in range(3):        try:            r = requests.get(API, params={"q": query}, timeout=8)            if r.status_code == 429:                time.sleep(2 ** attempt)      # 1s, 2s, 4s                continue            r.raise_for_status()            hits = r.json().get("results", [])[:5]            if not hits:                return (f"No results for {query!r}. Try broader terms, "                        f"or state that the information is unavailable.")            return "\n".join(f"- {h['title']}: {h['snippet']}" for h in hits)        except requests.Timeout:            if attempt == 2:                return "ERROR: search timed out three times. Proceed without it."        except requests.HTTPError as e:            if e.response.status_code in (401, 403):                raise                          # fatal: bad credentials            return f"ERROR: search failed ({e.response.status_code})."    return "ERROR: search unavailable."

Notice the empty-results branch. Returning "" for no results is what caused the forty-one-search loop in the opening story: an empty string looks to the model like "something went slightly wrong, try again". A sentence that explicitly says "or state that the information is unavailable" gives it a legitimate exit.

Decide, for every tool, whether its failure is the model's problem or yours. Return a string when the model can fix it; raise only when no amount of cleverness can.

Debugging: seeing what the model saw

Streaming the steps

Python
question = {"messages": [{"role": "user",                          "content": "Compare Tesla and Ford Q3 revenue"}]}for chunk in agent.stream(question, stream_mode="updates"):    for node, update in chunk.items():        for m in (update or {}).get("messages", []):            if m.type == "ai" and m.tool_calls:                for c in m.tool_calls:                    print(f"CALL   {c['name']}({c['args']})")            elif m.type == "tool":                print(f"RESULT {m.content[:160]}")            else:                print(f"ANSWER {m.content}")

Callbacks for structured traces

Python
from langchain_core.callbacks import BaseCallbackHandlerimport time, json, uuidclass TraceHandler(BaseCallbackHandler):    def __init__(self):        self.run_id = str(uuid.uuid4()); self.t0 = {}; self.calls = 0    def on_tool_start(self, serialized, input_str, **kw):        name = serialized.get("name", "?")        self.t0[name] = time.time(); self.calls += 1    def on_tool_end(self, output, **kw):        name = kw.get("name", "?")        print(json.dumps({"run": self.run_id, "tool": name,                          "ms": round((time.time() - self.t0.get(name, 0)) * 1000),                          "chars": len(str(output))}))    def on_llm_end(self, response, **kw):        u = response.generations[0][0].message.usage_metadata        print(json.dumps({"run": self.run_id, "usage": u}))agent.invoke({"messages": [{"role": "user", "content": "..."}]},             config={"callbacks": [TraceHandler()]})

Setting the environment variables LANGSMITH_TRACING=true and LANGSMITH_API_KEY=... sends the same information to a hosted trace viewer, which is worth it once traces get long enough that reading JSON in a terminal stops working.

The four ways LangChain agents go wrong

Vague tool descriptions

Covered above, and it remains the single largest lever. A useful discipline: write the description, then read it as if you were a contractor who has never seen the codebase. If you could not tell from that text alone when to call it and what to pass, the model cannot either.

Exceptions instead of error strings

A tool that raises on a 404 turns a recoverable situation into a dead request. Wrap the body, return a string starting with ERROR:, and reserve raising for genuinely fatal conditions.

Too many tools

Tool selection accuracy falls as the tool count rises, and the schemas cost tokens on every single step. Suppose each tool schema is about 180 tokens. With 5 tools that is 900 tokens per step; over a 6-step run, 5,400 tokens. With 25 tools it is 4,500 per step and 27,000 per run — an extra 0.065 dollars per request at an example rate of 3 dollars per million tokens, before the model has done anything useful. Worse, the model now has 25 plausible options and picks wrong more often.

The fix is routing: a cheap first classification picks a subset of tools, and the agent is constructed with only those.

Python
TOOLSETS = {    "finance":  [get_stock_price, get_financials, calculate],    "research": [search_web, read_url, summarise],    "internal": [query_db, lookup_order, calculate],}def route(question: str) -> str:    label = classifier_llm.invoke(        f"Reply with exactly one word - finance, research or internal:\n{question}"    ).content.strip().lower()    return label if label in TOOLSETS else "research"tools = TOOLSETS[route(user_question)]     # 3 tools in context, not 25

Unbounded tool use

Always set a model-call limit and a wall-clock deadline. Without them the only stop is LangGraph's recursion limit, which by default is in the thousands of steps. A cap of 6 model calls with a 45-second deadline bounds a runaway request to roughly 6 model calls instead of 41, and turns a failure costing tens of dollars into one costing well under a dollar. create_agent has no deadline argument, so put one around the call — await asyncio.wait_for(agent.ainvoke(...), timeout=45). Then handle the cap explicitly rather than shipping the terse "Model call limits exceeded" message to users:

Python
out = agent.invoke({"messages": [{"role": "user", "content": q}]})model_turns = sum(m.type == "ai" for m in out["messages"])answer = out["messages"][-1].contentif model_turns >= 6:    log.warning("hit model-call cap", extra={"q": q, "tools":                [name for name, _ in tool_steps(out["messages"])]})    found = "\n".join(f"- {r[:200]}" for _, r in tool_steps(out["messages"]))    answer = ("I could not finish this within my step budget. "              "Here is what I found so far:\n" + found)

Making this survive contact with users

When you put a LangChain agent in front of real traffic, three habits separate the systems that hold up from the ones that generate support tickets.

Treat tool descriptions as versioned artefacts. They are prompt engineering, they change behaviour, and a one-word edit can change your tool-selection rate by tens of percentage points. Keep them in code review. When an agent misbehaves, diff the description before you touch anything else.

Build a fixed evaluation set of twenty questions with known correct tool sequences. For each, record which tools should be called. Run it after every change to a prompt, a description or a model version. Score two things separately: did it call the right tools, and was the final answer right. They fail independently, and knowing which one broke tells you where to look.

Assume every number in a final answer is a hallucination until a tool result proves otherwise. The opening bug — a confident "2.1 million dollars" from nowhere — is the characteristic failure of tool-using agents, and it is invisible unless you check. Keep the tool results — the ToolMessages in the returned messages — and for numeric answers, verify programmatically that each figure appears in some tool result. If it does not, do not show it to the user.