Course Content
AI Agent Frameworks
4 sections · 15 lessons
Exercise 3: Compare Tool-Calling Patterns
Two engineers are given the same specification: build a document analyser that extracts key points, counts words, assesses sentiment, and produces a summary. One builds it as a tool-calling agent. The other builds it as a graph. Both work. In the demo they look identical.
Then the requirements change, as they always do. The product manager asks for three things: "can the summary always mention the sentiment?", "can we retry if the summary is under 40 words?", and "can we skip sentiment for internal documents?"
The graph author adds a node, a conditional edge and a counter in about twenty minutes. The agent author rewrites the system prompt, tests it, discovers the agent honours the new rule roughly seven times out of ten, adds an emphatic sentence, gets to eight, and spends the afternoon there.
That is the comparison this exercise is designed to make you feel rather than read. You will build the same analyser twice, run both against the same three documents, record the same measurements, and reach your own conclusion about when each shape is right.
The specification
Input: a document as a string. Output: a report containing key points, word count, sentiment, and a summary. Four capabilities, and they are deliberately chosen to have a mix of properties.
| Capability | Deterministic? | Depends on |
|---|---|---|
count_words | Yes — pure Python | Nothing |
extract_key_points | No — model call | The document |
analyse_sentiment | No — model call | The document |
summarise | No — model call | Key points and sentiment |
The last row is the crux of the whole exercise. summarise has a genuine dependency: it needs the other two to have run first. In a graph that dependency is an edge. In an agent it is a sentence in a prompt, and the difference between an edge and a sentence is the difference between a guarantee and a tendency.
Part 1 — the agent implementation
1import os2from langchain_core.tools import tool3from langchain_anthropic import ChatAnthropic4from langchain.agents import create_agent5from langchain.agents.middleware import ModelCallLimitMiddleware67llm = ChatAnthropic(model=os.getenv("AGENT_MODEL", "claude-sonnet-5"))89@tool10def count_words(text: str) -> str:11 """Count words, sentences and characters in a document.12 Pass the full document text. Returns exact counts."""13 words = text.split()14 sentences = [s for s in text.replace("!", ".").replace("?", ".").split(".")15 if s.strip()]16 return (f"words={len(words)} sentences={len(sentences)} "17 f"chars={len(text)} avg_words_per_sentence="18 f"{len(words) / max(len(sentences), 1):.1f}")1920@tool21def extract_key_points(text: str) -> str:22 """Extract the 3-5 most important claims from a document as a bullet list.23 Use before summarising. Pass the full document text."""24 return llm.invoke("List the 3-5 most important claims, one per line, "25 "as '- claim'. No preamble.\n\n" + text).content2627@tool28def analyse_sentiment(text: str) -> str:29 """Judge the overall tone of a document.30 Returns one of positive / negative / neutral / mixed, plus a confidence31 from 0 to 1 and a one-line justification."""32 return llm.invoke("Reply exactly as: LABEL | CONFIDENCE | REASON\n"33 "LABEL is positive, negative, neutral or mixed.\n\n"34 + text).content3536@tool37def summarise(key_points: str, sentiment: str) -> str:38 """Write a 60-80 word summary. REQUIRES the outputs of extract_key_points39 and analyse_sentiment - call both of those first and pass their results40 here verbatim. The summary must mention the sentiment explicitly."""41 return llm.invoke(f"Write a 60-80 word summary that explicitly names the "42 f"tone.\n\nKey points:\n{key_points}\n\n"43 f"Sentiment: {sentiment}").content4445TOOLS = [count_words, extract_key_points, analyse_sentiment, summarise]4647agent_app = create_agent(48 model=llm, tools=TOOLS,49 system_prompt=(50 "Analyse the document the user provides. You MUST call, in this order:\n"51 "1. count_words\n2. extract_key_points\n3. analyse_sentiment\n"52 "4. summarise, passing the outputs of steps 2 and 3.\n"53 "Then present a report with headings: Key Points, Statistics, "54 "Sentiment, Summary. Do not skip any step."),55 middleware=[ModelCallLimitMiddleware(run_limit=8, exit_behavior="end")],56)Read the system prompt carefully. It is a numbered procedure written in English, addressed to a model that is free to ignore it. That freedom is the agent's defining property — sometimes an asset, here a liability.
Part 2 — the graph implementation
1from typing import TypedDict, Annotated, Literal2import operator3from langgraph.graph import StateGraph, START, END45class DocState(TypedDict):6 document: str7 stats: str8 key_points: str9 sentiment: str10 summary: str11 attempts: int12 trace: Annotated[list[str], operator.add]1314def stats_node(s):15 words = s["document"].split()16 return {"stats": f"words={len(words)} chars={len(s['document'])}",17 "attempts": 0, "trace": ["stats"]}1819def points_node(s):20 return {"key_points": llm.invoke(21 "List the 3-5 most important claims, one per line, as "22 "'- claim'. No preamble.\n\n" + s["document"]).content,23 "trace": ["points"]}2425def sentiment_node(s):26 return {"sentiment": llm.invoke(27 "Reply exactly as: LABEL | CONFIDENCE | REASON\n"28 "LABEL is positive, negative, neutral or mixed.\n\n"29 + s["document"]).content,30 "trace": ["sentiment"]}3132def summary_node(s):33 hint = "" if s["attempts"] == 0 else "\nYour last attempt was too short."34 return {"summary": llm.invoke(35 f"Write a 60-80 word summary that explicitly names the tone."36 f"{hint}\n\nKey points:\n{s['key_points']}\n\n"37 f"Sentiment: {s['sentiment']}").content,38 "attempts": s["attempts"] + 1,39 "trace": [f"summary attempt {s['attempts'] + 1}"]}4041def check(s) -> Literal["ok", "redo"]:42 long_enough = len(s["summary"].split()) >= 4043 names_tone = any(w in s["summary"].lower()44 for w in ("positive", "negative", "neutral", "mixed"))45 if (long_enough and names_tone) or s["attempts"] >= 2:46 return "ok"47 return "redo"4849g = StateGraph(DocState)50for n, f in [("stats", stats_node), ("points", points_node),51 ("sentiment", sentiment_node), ("summary", summary_node)]:52 g.add_node(n, f)5354g.add_edge(START, "stats")55g.add_edge("stats", "points") # fan out56g.add_edge("stats", "sentiment") # these two run in parallel57g.add_edge("points", "summary") # fan in - summary waits for both58g.add_edge("sentiment", "summary")59g.add_conditional_edges("summary", check, {"ok": END, "redo": "summary"})6061graph_app = g.compile()Three things in that wiring are impossible to express in the agent version:
- Parallelism.
pointsandsentimentboth depend onstatsand on nothing else, so LangGraph runs them concurrently. - A hard dependency.
summaryhas incoming edges from both, so it cannot start until both have finished. It is not asked to wait; it is unable to run. - A checked retry.
checkis Python. "At least 40 words and names the tone" is measured, not judged, and the retry happens at most twice.
In an agent, "call these in order" is a request. In a graph, it is a topology. Requests are honoured most of the time; topologies are honoured every time.
Part 3 — measure them
Run both against three documents chosen to stress different things: a positive product announcement of about 200 words, a critical review of about 400 words, and a neutral technical spec of about 800 words.
1import time, json23def measure(name, fn, doc):4 t0 = time.perf_counter()5 out = fn(doc)6 return {"impl": name, "seconds": round(time.perf_counter() - t0, 2), **out}78def run_agent(doc):9 r = agent_app.invoke({"messages": [{"role": "user", "content": doc}]})10 tools = [m.name for m in r["messages"] if m.type == "tool"]11 return {"calls": len(tools), "order": tools,12 "summary_words": len(r["messages"][-1].content.split())}1314def run_graph(doc):15 r = graph_app.invoke({"document": doc})16 return {"calls": len(r["trace"]), "order": r["trace"],17 "summary_words": len(r["summary"].split())}1819for doc_name, doc in DOCS.items():20 for name, fn in (("agent", run_agent), ("graph", run_graph)):21 print(json.dumps({"doc": doc_name, **measure(name, fn, doc)}))Run each combination five times. Variation between runs is itself a result — and it is the result that matters most.
What a typical set of numbers looks like
| Measure | Agent | Graph | Why |
|---|---|---|---|
| Lines of code | ~55 | ~70 | Graph pays for explicit wiring |
| Model calls per run | 5 agent turns + 3 model-backed tools = 8 | 3 | The agent needs a turn to decide each call |
| Wall-clock, 800-word doc | ~14 s | ~8 s | Parallel branches, no decision turns |
| Correct step order, 15 runs | 13/15 | 15/15 | Topology versus instruction |
| Summary meets the 40-word rule | 11/15 | 15/15 | The graph checks and retries |
| Handles an unexpected request | Yes — improvises | No — runs the fixed path | Freedom cuts both ways |
Do the cost arithmetic yourself, because it is the least intuitive column. The agent makes eight model calls (four turns that each choose a tool, one final turn, and three tools that call the model themselves; count_words is plain Python). The graph makes three, one each for key points, sentiment and summary. Suppose the document plus the accumulated messages averages 2,000 input tokens per call, at an example rate of 3 dollars per million:
- Agent: 8×2,000=16,000 input tokens ≈ 0.048 dollars per document
- Graph: 3×2,000=6,000 input tokens ≈ 0.018 dollars per document
About a 2.7x difference. Across 100,000 documents a month that is 4,800 dollars versus 1,800 dollars. The gap comes almost entirely from the agent's decision turns — the model calls whose only output is "now call analyse_sentiment". A graph has no decision turns because the decisions were made at design time.
The comparison to fill in
Record your own numbers in this shape. The empty cells are the exercise.
| Dimension | Agent | Graph | Winner and why |
|---|---|---|---|
| Setup time (first working version) | |||
| Median latency, 800-word doc | |||
| Model calls per run | |||
| Step-order compliance, 15 runs | |||
| Time to add "skip sentiment for internal docs" | |||
| Time to diagnose a wrong summary | |||
| Behaviour when a tool raises |
The three change requests
Now implement the product manager's three requests in both versions and time yourself. This is where the comparison stops being academic.
| Request | Agent change | Graph change |
|---|---|---|
| Summary must always mention sentiment | Strengthen the prompt; verify empirically; accept a failure rate | Add the condition to check; failure rate becomes zero |
| Retry if the summary is under 40 words | No reliable mechanism — the model must notice and re-call itself | Already there: one clause in check |
| Skip sentiment for internal documents | Add a condition in prose and hope | Conditional edge out of stats |
Notice that the second request has no good agent answer at all. You can tell the model "check your summary is at least 40 words and rewrite it if not", and it will comply most of the time — but "most of the time" is not a retry mechanism, it is a probability. This is the clearest case in the exercise where the two shapes are not merely different in convenience but different in what they can promise.
A decision the model makes at run time is a decision you pay a model call for. If you already knew the answer at design time, you are buying it twice.
The failure modes you should expect to see
| Observation | What it means |
|---|---|
Agent calls summarise before analyse_sentiment | Ordering by prose is advisory; the model judged sentiment unnecessary |
Agent passes an empty string as sentiment | It satisfied the schema without satisfying the intent — schemas check shape, not meaning |
Agent skips count_words and estimates | The tool's value was not obvious from its description |
Graph's summary node sees an empty key_points | An edge is missing, so the node ran before its dependency |
Graph trace shows one entry | Missing operator.add reducer |
| Graph never retries | attempts reset somewhere, or check too lenient |
Choosing between them for real work
The conclusion this exercise should push you towards is not "graphs are better". It is that the two shapes answer a different question, and you should pick by asking which question your problem actually poses.
If you can write the steps down in order before you see the input, you have a workflow, and a workflow expressed as a graph gives you compliance, parallelism and cheaper runs. The document analyser is exactly this: four steps, one dependency, known in advance. Building it as an agent means paying five extra model calls per document for a decision you had already made.
If you genuinely cannot write the steps down — because the right sequence depends on what the input turns out to contain, and there are ten plausible paths — the agent's improvisation is the whole product. Forcing that into a graph means enumerating branches you cannot enumerate.
The most useful pattern in practice is neither pure: a graph for the skeleton, with one node that runs a tool-calling agent for the genuinely open-ended step. The analyser could route documents to a graph path when they match a known type and to an agent when they do not. You get guarantees where you can specify them and flexibility where you cannot, and — importantly — the boundary between the two is a node you can point at, rather than a hope distributed through a system prompt.
One last practical note from doing this comparison honestly: run each version fifteen times, not once. Both implementations will look correct on a single run. The difference lives entirely in the variance, and a single run is precisely the measurement that hides it.