AI Agent Frameworks

Exercise 3: Compare Tool-Calling Patterns


Two engineers are given the same specification: build a document analyser that extracts key points, counts words, assesses sentiment, and produces a summary. One builds it as a tool-calling agent. The other builds it as a graph. Both work. In the demo they look identical.

Then the requirements change, as they always do. The product manager asks for three things: "can the summary always mention the sentiment?", "can we retry if the summary is under 40 words?", and "can we skip sentiment for internal documents?"

The graph author adds a node, a conditional edge and a counter in about twenty minutes. The agent author rewrites the system prompt, tests it, discovers the agent honours the new rule roughly seven times out of ten, adds an emphatic sentence, gets to eight, and spends the afternoon there.

That is the comparison this exercise is designed to make you feel rather than read. You will build the same analyser twice, run both against the same three documents, record the same measurements, and reach your own conclusion about when each shape is right.

The same analyser, two shapesTool-calling agent• Model decides the order at run time• Fewer lines, faster to a demo• Step count varies run to run• A change request is a prompt editExplicit graph• You fix the order at build time• More wiring, deterministic path• Latency and cost are predictable• A change request is a new node
They look identical in the demo; the difference only appears in the variance and the third change request.

The specification

Input: a document as a string. Output: a report containing key points, word count, sentiment, and a summary. Four capabilities, and they are deliberately chosen to have a mix of properties.

CapabilityDeterministic?Depends on
count_wordsYes — pure PythonNothing
extract_key_pointsNo — model callThe document
analyse_sentimentNo — model callThe document
summariseNo — model callKey points and sentiment

The last row is the crux of the whole exercise. summarise has a genuine dependency: it needs the other two to have run first. In a graph that dependency is an edge. In an agent it is a sentence in a prompt, and the difference between an edge and a sentence is the difference between a guarantee and a tendency.

Part 1 — the agent implementation

Python
import osfrom langchain_core.tools import toolfrom langchain_anthropic import ChatAnthropicfrom langchain.agents import create_agentfrom langchain.agents.middleware import ModelCallLimitMiddlewarellm = ChatAnthropic(model=os.getenv("AGENT_MODEL", "claude-sonnet-5"))@tooldef count_words(text: str) -> str:    """Count words, sentences and characters in a document.    Pass the full document text. Returns exact counts."""    words = text.split()    sentences = [s for s in text.replace("!", ".").replace("?", ".").split(".")                 if s.strip()]    return (f"words={len(words)} sentences={len(sentences)} "            f"chars={len(text)} avg_words_per_sentence="            f"{len(words) / max(len(sentences), 1):.1f}")@tooldef extract_key_points(text: str) -> str:    """Extract the 3-5 most important claims from a document as a bullet list.    Use before summarising. Pass the full document text."""    return llm.invoke("List the 3-5 most important claims, one per line, "                      "as '- claim'. No preamble.\n\n" + text).content@tooldef analyse_sentiment(text: str) -> str:    """Judge the overall tone of a document.    Returns one of positive / negative / neutral / mixed, plus a confidence    from 0 to 1 and a one-line justification."""    return llm.invoke("Reply exactly as: LABEL | CONFIDENCE | REASON\n"                      "LABEL is positive, negative, neutral or mixed.\n\n"                      + text).content@tooldef summarise(key_points: str, sentiment: str) -> str:    """Write a 60-80 word summary. REQUIRES the outputs of extract_key_points    and analyse_sentiment - call both of those first and pass their results    here verbatim. The summary must mention the sentiment explicitly."""    return llm.invoke(f"Write a 60-80 word summary that explicitly names the "                      f"tone.\n\nKey points:\n{key_points}\n\n"                      f"Sentiment: {sentiment}").contentTOOLS = [count_words, extract_key_points, analyse_sentiment, summarise]agent_app = create_agent(    model=llm, tools=TOOLS,    system_prompt=(        "Analyse the document the user provides. You MUST call, in this order:\n"        "1. count_words\n2. extract_key_points\n3. analyse_sentiment\n"        "4. summarise, passing the outputs of steps 2 and 3.\n"        "Then present a report with headings: Key Points, Statistics, "        "Sentiment, Summary. Do not skip any step."),    middleware=[ModelCallLimitMiddleware(run_limit=8, exit_behavior="end")],)

Read the system prompt carefully. It is a numbered procedure written in English, addressed to a model that is free to ignore it. That freedom is the agent's defining property — sometimes an asset, here a liability.

Part 2 — the graph implementation

Python
from typing import TypedDict, Annotated, Literalimport operatorfrom langgraph.graph import StateGraph, START, ENDclass DocState(TypedDict):    document: str    stats: str    key_points: str    sentiment: str    summary: str    attempts: int    trace: Annotated[list[str], operator.add]def stats_node(s):    words = s["document"].split()    return {"stats": f"words={len(words)} chars={len(s['document'])}",            "attempts": 0, "trace": ["stats"]}def points_node(s):    return {"key_points": llm.invoke(                "List the 3-5 most important claims, one per line, as "                "'- claim'. No preamble.\n\n" + s["document"]).content,            "trace": ["points"]}def sentiment_node(s):    return {"sentiment": llm.invoke(                "Reply exactly as: LABEL | CONFIDENCE | REASON\n"                "LABEL is positive, negative, neutral or mixed.\n\n"                + s["document"]).content,            "trace": ["sentiment"]}def summary_node(s):    hint = "" if s["attempts"] == 0 else "\nYour last attempt was too short."    return {"summary": llm.invoke(                f"Write a 60-80 word summary that explicitly names the tone."                f"{hint}\n\nKey points:\n{s['key_points']}\n\n"                f"Sentiment: {s['sentiment']}").content,            "attempts": s["attempts"] + 1,            "trace": [f"summary attempt {s['attempts'] + 1}"]}def check(s) -> Literal["ok", "redo"]:    long_enough = len(s["summary"].split()) >= 40    names_tone = any(w in s["summary"].lower()                     for w in ("positive", "negative", "neutral", "mixed"))    if (long_enough and names_tone) or s["attempts"] >= 2:        return "ok"    return "redo"g = StateGraph(DocState)for n, f in [("stats", stats_node), ("points", points_node),             ("sentiment", sentiment_node), ("summary", summary_node)]:    g.add_node(n, f)g.add_edge(START, "stats")g.add_edge("stats", "points")        # fan outg.add_edge("stats", "sentiment")     # these two run in parallelg.add_edge("points", "summary")      # fan in - summary waits for bothg.add_edge("sentiment", "summary")g.add_conditional_edges("summary", check, {"ok": END, "redo": "summary"})graph_app = g.compile()

Three things in that wiring are impossible to express in the agent version:

  • Parallelism. points and sentiment both depend on stats and on nothing else, so LangGraph runs them concurrently.
  • A hard dependency. summary has incoming edges from both, so it cannot start until both have finished. It is not asked to wait; it is unable to run.
  • A checked retry. check is Python. "At least 40 words and names the tone" is measured, not judged, and the retry happens at most twice.

In an agent, "call these in order" is a request. In a graph, it is a topology. Requests are honoured most of the time; topologies are honoured every time.

Part 3 — measure them

Run both against three documents chosen to stress different things: a positive product announcement of about 200 words, a critical review of about 400 words, and a neutral technical spec of about 800 words.

Python
import time, jsondef measure(name, fn, doc):    t0 = time.perf_counter()    out = fn(doc)    return {"impl": name, "seconds": round(time.perf_counter() - t0, 2), **out}def run_agent(doc):    r = agent_app.invoke({"messages": [{"role": "user", "content": doc}]})    tools = [m.name for m in r["messages"] if m.type == "tool"]    return {"calls": len(tools), "order": tools,            "summary_words": len(r["messages"][-1].content.split())}def run_graph(doc):    r = graph_app.invoke({"document": doc})    return {"calls": len(r["trace"]), "order": r["trace"],            "summary_words": len(r["summary"].split())}for doc_name, doc in DOCS.items():    for name, fn in (("agent", run_agent), ("graph", run_graph)):        print(json.dumps({"doc": doc_name, **measure(name, fn, doc)}))

Run each combination five times. Variation between runs is itself a result — and it is the result that matters most.

What a typical set of numbers looks like

MeasureAgentGraphWhy
Lines of code~55~70Graph pays for explicit wiring
Model calls per run5 agent turns + 3 model-backed tools = 83The agent needs a turn to decide each call
Wall-clock, 800-word doc~14 s~8 sParallel branches, no decision turns
Correct step order, 15 runs13/1515/15Topology versus instruction
Summary meets the 40-word rule11/1515/15The graph checks and retries
Handles an unexpected requestYes — improvisesNo — runs the fixed pathFreedom cuts both ways

Do the cost arithmetic yourself, because it is the least intuitive column. The agent makes eight model calls (four turns that each choose a tool, one final turn, and three tools that call the model themselves; count_words is plain Python). The graph makes three, one each for key points, sentiment and summary. Suppose the document plus the accumulated messages averages 2,000 input tokens per call, at an example rate of 3 dollars per million:

  • Agent: 8×2,000=16,0008 \times 2{,}000 = 16{,}000 input tokens ≈ 0.048 dollars per document
  • Graph: 3×2,000=6,0003 \times 2{,}000 = 6{,}000 input tokens ≈ 0.018 dollars per document

About a 2.7x difference. Across 100,000 documents a month that is 4,800 dollars versus 1,800 dollars. The gap comes almost entirely from the agent's decision turns — the model calls whose only output is "now call analyse_sentiment". A graph has no decision turns because the decisions were made at design time.

The comparison to fill in

Record your own numbers in this shape. The empty cells are the exercise.

DimensionAgentGraphWinner and why
Setup time (first working version)
Median latency, 800-word doc
Model calls per run
Step-order compliance, 15 runs
Time to add "skip sentiment for internal docs"
Time to diagnose a wrong summary
Behaviour when a tool raises

The three change requests

Now implement the product manager's three requests in both versions and time yourself. This is where the comparison stops being academic.

RequestAgent changeGraph change
Summary must always mention sentimentStrengthen the prompt; verify empirically; accept a failure rateAdd the condition to check; failure rate becomes zero
Retry if the summary is under 40 wordsNo reliable mechanism — the model must notice and re-call itselfAlready there: one clause in check
Skip sentiment for internal documentsAdd a condition in prose and hopeConditional edge out of stats

Notice that the second request has no good agent answer at all. You can tell the model "check your summary is at least 40 words and rewrite it if not", and it will comply most of the time — but "most of the time" is not a retry mechanism, it is a probability. This is the clearest case in the exercise where the two shapes are not merely different in convenience but different in what they can promise.

A decision the model makes at run time is a decision you pay a model call for. If you already knew the answer at design time, you are buying it twice.

The failure modes you should expect to see

ObservationWhat it means
Agent calls summarise before analyse_sentimentOrdering by prose is advisory; the model judged sentiment unnecessary
Agent passes an empty string as sentimentIt satisfied the schema without satisfying the intent — schemas check shape, not meaning
Agent skips count_words and estimatesThe tool's value was not obvious from its description
Graph's summary node sees an empty key_pointsAn edge is missing, so the node ran before its dependency
Graph trace shows one entryMissing operator.add reducer
Graph never retriesattempts reset somewhere, or check too lenient

Choosing between them for real work

The conclusion this exercise should push you towards is not "graphs are better". It is that the two shapes answer a different question, and you should pick by asking which question your problem actually poses.

If you can write the steps down in order before you see the input, you have a workflow, and a workflow expressed as a graph gives you compliance, parallelism and cheaper runs. The document analyser is exactly this: four steps, one dependency, known in advance. Building it as an agent means paying five extra model calls per document for a decision you had already made.

If you genuinely cannot write the steps down — because the right sequence depends on what the input turns out to contain, and there are ten plausible paths — the agent's improvisation is the whole product. Forcing that into a graph means enumerating branches you cannot enumerate.

The most useful pattern in practice is neither pure: a graph for the skeleton, with one node that runs a tool-calling agent for the genuinely open-ended step. The analyser could route documents to a graph path when they match a known type and to an agent when they do not. You get guarantees where you can specify them and flexibility where you cannot, and — importantly — the boundary between the two is a node you can point at, rather than a hope distributed through a system prompt.

One last practical note from doing this comparison honestly: run each version fifteen times, not once. Both implementations will look correct on a single run. The difference lives entirely in the variance, and a single run is precisely the measurement that hides it.