AI Agent Frameworks

Exercise 2: Create a LangGraph Workflow


Here is a support desk with a real problem. Three kinds of question arrive: "what's 18% VAT on 340?", "who's the current CFO of Siemens?", and "why does this Python function return None?". They need completely different handling — an exact calculator, a live web search, and a code-reading model with a long context. A single agent with all three tools loaded answers all three, but it loads roughly 540 tokens of tool schemas into every request including the trivial ones, and it picks the wrong tool about one time in six.

You are going to build the alternative: a graph that classifies first, then routes. One cheap model call decides which of three specialised branches handles the request. Each branch sees only its own tools. A shared formatting node produces the final answer. And a validation step sends bad output back for one retry, with a hard limit so it cannot loop.

By the end you will have written a state schema with reducers, three node functions, one conditional edge, one cycle with a counter, and a compiled graph you can invoke. That is the complete LangGraph vocabulary — everything larger is the same pieces repeated.

Classify once, then dispatchclassifier nodearithmetic?lookup or code?calculatorweb searchcode model
Routing on a classifier keeps each branch's prompt small — one shared prompt would have to serve all three questions.

Setup

Bash
pip install langgraph langchain-anthropic python-dotenv
Python
import os, ast, operatorfrom typing import TypedDict, Annotated, Literalfrom dotenv import load_dotenvfrom langchain_anthropic import ChatAnthropicfrom langgraph.graph import StateGraph, START, ENDload_dotenv()fast  = ChatAnthropic(model=os.getenv("FAST_MODEL", "claude-haiku-4-5"))smart = ChatAnthropic(model=os.getenv("SMART_MODEL", "claude-sonnet-5"))

Two models, deliberately. Classification is a one-word decision that a small fast model does as well as a large one, faster and at a fraction of the price (at the time of writing, about half the per-token price of the larger model). Answering is where you spend.

Requirement 1 — the state schema

State is a TypedDict. Every node receives the whole thing and returns a partial update, which LangGraph merges in. The merge rule per field is what you must decide up front.

Python
class State(TypedDict):    question: str                              # set once, overwrite    route: str                                 # "math" | "search" | "code"    answer: str                                # overwritten each attempt    verdict: str                               # "PASS" or a complaint    attempts: int                              # overwritten, incremented by node    trace: Annotated[list[str], operator.add]  # ACCUMULATES

Only one field has a reducer, and it is the one you want to grow. The default behaviour for the others is overwrite, which is exactly right for them.

Get this wrong and the symptom is silent. Suppose trace had no reducer and four nodes each returned {"trace": ["classified as math"]}, {"trace": ["computed 61.2"]} and so on. The final state holds one entry, not four, and there is no error to tell you.

FieldWritten byWanted behaviourAnnotation
questionThe callerSet oncenone
routeclassifyLatest winsnone
answerThree branch nodesLatest wins — a retry replaces the old answernone
attemptsBranch nodesLatest wins (node computes old + 1)none
traceEvery nodeGrowoperator.add

Before writing a single node, go through the state schema field by field and answer one question for each: if two nodes write this, should the second replace the first or add to it? That decision is the reducer, and it is unrecoverable once the run has finished.

Requirement 2 — the classifier node

A node is a plain function. Take state, return a dict of updates. Nothing more.

Python
VALID = {"math", "search", "code"}def classify(state: State) -> dict:    raw = fast.invoke(        "Classify the user's question into exactly one category.\n"        "math   - arithmetic, percentages, unit conversion, numeric comparison\n"        "search - facts that may have changed: news, prices, people's roles\n"        "code   - reading, explaining, debugging or writing source code\n"        "Reply with the single word and nothing else.\n\n"        f"Question: {state['question']}"    ).content.strip().lower()    route = raw if raw in VALID else "search"    return {"route": route,            "attempts": 0,            "trace": [f"classify: raw={raw!r} -> {route}"]}

The line route = raw if raw in VALID else "search" is the important one. Models occasionally reply "Math." or "This is a math question", and an unrecognised label passed into a conditional edge raises a KeyError that kills the run. Validating against a set and falling back to the safest branch converts a crash into a slightly suboptimal route.

Choose your fallback deliberately. search is the right default here because it is the most general branch — it can produce a reasonable answer to a math or code question, whereas the calculator branch cannot answer "who is the CFO of Siemens".

Requirement 3 — the three branches

Python
_OPS = {ast.Add: operator.add, ast.Sub: operator.sub, ast.Mult: operator.mul,        ast.Div: operator.truediv, ast.Pow: operator.pow, ast.USub: operator.neg}def _calc(node):    if isinstance(node, ast.Constant) and isinstance(node.value, (int, float)):        return node.value    if isinstance(node, ast.BinOp) and type(node.op) in _OPS:        return _OPS[type(node.op)](_calc(node.left), _calc(node.right))    if isinstance(node, ast.UnaryOp) and type(node.op) in _OPS:        return _OPS[type(node.op)](_calc(node.operand))    raise ValueError("unsupported expression")def math_node(state: State) -> dict:    expr = fast.invoke(        "Rewrite this as a single arithmetic expression using only digits and "        "+ - * / ( ) ** . No words, no units, no percent signs. "        "Write percentages as decimals.\n\n" + state["question"]    ).content.strip().strip("`")    try:        value = _calc(ast.parse(expr, mode="eval").body)        answer = f"{state['question'].rstrip('?')} = {value:,.4f}".rstrip("0").rstrip(".")    except Exception as e:        answer = f"Could not evaluate {expr!r}: {e}"    return {"answer": answer,            "attempts": state["attempts"] + 1,            "trace": [f"math: expr={expr!r}"]}def search_node(state: State) -> dict:    from ddgs import DDGS    with DDGS() as d:        hits = list(d.text(state["question"], max_results=4))    context = "\n".join(f"- {h['title']}: {h['body'][:200]}" for h in hits) or "none"    answer = smart.invoke(        "Answer the question using ONLY these snippets. If they do not contain "        "the answer, say so plainly.\n\n"        f"Snippets:\n{context}\n\nQuestion: {state['question']}"    ).content    return {"answer": answer,            "attempts": state["attempts"] + 1,            "trace": [f"search: {len(hits)} hits"]}def code_node(state: State) -> dict:    answer = smart.invoke(        "You are a senior engineer. Answer precisely. If the question contains "        "code, quote the exact line that causes the behaviour before explaining "        "it.\n\n" + state["question"]    ).content    return {"answer": answer,            "attempts": state["attempts"] + 1,            "trace": ["code: answered"]}

Each branch does one thing and returns an answer. None of them decide what runs next — that is the graph's job, and keeping it out of the node functions is what makes them individually testable.

Requirement 4 — wiring the graph

Add a validation node and a cycle, so the workflow contains a genuine loop with a hard limit rather than an instruction that hopes to be obeyed.

Python
def validate(state: State) -> dict:    v = fast.invoke(        "Does this answer actually address the question? "        "Reply PASS, or one short sentence naming what is missing.\n\n"        f"Question: {state['question']}\nAnswer: {state['answer']}"    ).content.strip()    return {"verdict": v, "trace": [f"validate: {v[:60]}"]}def to_branch(state: State) -> Literal["math", "search", "code"]:    return state["route"]def after_validate(state: State) -> Literal["done", "retry"]:    if state["verdict"].upper().startswith("PASS"):        return "done"    if state["attempts"] >= 2:        return "done"          # give up gracefully rather than loop    return "retry"g = StateGraph(State)g.add_node("classify", classify)g.add_node("math", math_node)g.add_node("search", search_node)g.add_node("code", code_node)g.add_node("validate", validate)g.add_edge(START, "classify")g.add_conditional_edges("classify", to_branch,                        {"math": "math", "search": "search", "code": "code"})for b in ("math", "search", "code"):    g.add_edge(b, "validate")g.add_conditional_edges("validate", after_validate,                        {"done": END, "retry": "search"})app = g.compile()

The retry edge points at search rather than back at the original branch on purpose: if the first attempt failed validation, repeating the same branch usually reproduces the same failure, whereas a web search brings in information the first attempt did not have. That is a design choice, not a rule — send it back to to_branch instead if you prefer, but then the counter is doing all the work preventing an infinite loop.

The shape of the compiled graph:

Text
                    START                      |                  classify             /        |        \          math     search      code             \        |        /                  validate                   /     \             PASS /       \ FAIL and attempts < 2                 |         |                END      search  (loop back, max once)

A conditional edge is only as safe as its fallback. Any label the mapping does not contain is a crash, so validate the classifier's answer before the graph ever sees it.

Testing it

Python
CASES = [    ("What is 18% of 340?",                        "math"),    ("What is 847 * 293?",                         "math"),    ("Who is the current CEO of Siemens?",          "search"),    ("What happened to oil prices this week?",      "search"),    ("Why does this return None?\n\ndef f(x):\n    x.sort()\n    return x.sort()",                                                    "code"),    ("Explain Python's GIL in two sentences.",      "code"),]hits = 0for q, expected in CASES:    out = app.invoke({"question": q})    ok = out["route"] == expected    hits += ok    print(f"{'OK ' if ok else 'MISS'} route={out['route']:<7} expected={expected}")    for line in out["trace"]:        print("      " + line)print(f"routing accuracy: {hits}/{len(CASES)}")

Two of these have answers you can check by hand, which is the point of including them. 0.18×340=61.20.18 \times 340 = 61.2 and 847×293=248,171847 \times 293 = 248{,}171. If the math branch returns anything else, the expression-rewriting prompt is producing something the parser mangles — print expr from the trace and you will see it immediately, usually a stray % or a comma inside a number.

The third-from-last case is the interesting one. x.sort() sorts in place and returns None, so return x.sort() returns None — the code branch should quote that line. If it gets routed to search instead, your classifier prompt needs the word "debugging" made more prominent, because the model is latching onto "why does" as a question about the world.

Where this goes wrong

SymptomCauseFix
KeyError: 'Math.' from the conditional edgeClassifier returned a label not in the mappingValidate against a set, fall back to a default
trace has one entry at the endMissing Annotated[list, operator.add]Add the reducer
GraphRecursionErrorRetry edge with no attempt counter, or attempts reset inside the branchIncrement in the branch, check in after_validate
Node's change is invisible downstreamNode mutated state["x"] = ... instead of returningAlways return {...}
Everything routes to searchClassifier replying in a sentence, hitting the fallback every timePrint the raw reply; add "and nothing else" to the prompt
Retry never triggersValidator almost always says PASSMake the validator's criterion specific and checkable, not "is this good"

Worth extending

  • Add a fourth branch and measure the classifier again. Routing accuracy falls as categories multiply and their boundaries blur. Watching that happen with your own numbers is more convincing than being told.
  • Fan out instead of routing. For ambiguous questions, run search and code in parallel by adding both edges out of classify, then add a merge node. You will need a reducer on answer too — a list rather than a string.
  • Compile with a checkpointer (SqliteSaver, from langgraph-checkpoint-sqlite) and a thread_id. Kill the process mid-run and resume it. The state comes back.
  • Stream it: for ev in app.stream(inp, stream_mode="updates") gives you per-node progress, which is what you would show a user as "Classifying… Searching… Checking answer…".

What the graph bought you

Compare the finished system with the single agent described at the start. The agent carried all three tools in every request; the graph carries one branch's worth. On a math question that is roughly 180 tokens of schema instead of 540 — small per request, meaningful at 50,000 requests a month.

But token savings are the lesser prize. The real difference is that three properties of this workflow are now facts about the code rather than hopes about the model: a question is handled by exactly one branch, every answer passes through validation, and no request costs more than two attempts. You can read those guarantees off the graph definition and prove them with a unit test on after_validate that never calls a model at all.

That is the trade a graph asks you to make. You give up the model's freedom to improvise a sequence of steps, and you buy a workflow whose control flow you can point at, test and change. Take that trade when the steps are known and the guarantees matter. Refuse it when the whole value of the system is that you did not know in advance what it would need to do.