Course Content
AI Agent Frameworks
4 sections · 15 lessons
Exercise 2: Create a LangGraph Workflow
Here is a support desk with a real problem. Three kinds of question arrive: "what's 18% VAT on 340?", "who's the current CFO of Siemens?", and "why does this Python function return None?". They need completely different handling — an exact calculator, a live web search, and a code-reading model with a long context. A single agent with all three tools loaded answers all three, but it loads roughly 540 tokens of tool schemas into every request including the trivial ones, and it picks the wrong tool about one time in six.
You are going to build the alternative: a graph that classifies first, then routes. One cheap model call decides which of three specialised branches handles the request. Each branch sees only its own tools. A shared formatting node produces the final answer. And a validation step sends bad output back for one retry, with a hard limit so it cannot loop.
By the end you will have written a state schema with reducers, three node functions, one conditional edge, one cycle with a counter, and a compiled graph you can invoke. That is the complete LangGraph vocabulary — everything larger is the same pieces repeated.
Setup
pip install langgraph langchain-anthropic python-dotenv1import os, ast, operator2from typing import TypedDict, Annotated, Literal3from dotenv import load_dotenv4from langchain_anthropic import ChatAnthropic5from langgraph.graph import StateGraph, START, END67load_dotenv()8fast = ChatAnthropic(model=os.getenv("FAST_MODEL", "claude-haiku-4-5"))9smart = ChatAnthropic(model=os.getenv("SMART_MODEL", "claude-sonnet-5"))Two models, deliberately. Classification is a one-word decision that a small fast model does as well as a large one, faster and at a fraction of the price (at the time of writing, about half the per-token price of the larger model). Answering is where you spend.
Requirement 1 — the state schema
State is a TypedDict. Every node receives the whole thing and returns a partial update, which LangGraph merges in. The merge rule per field is what you must decide up front.
1class State(TypedDict):2 question: str # set once, overwrite3 route: str # "math" | "search" | "code"4 answer: str # overwritten each attempt5 verdict: str # "PASS" or a complaint6 attempts: int # overwritten, incremented by node7 trace: Annotated[list[str], operator.add] # ACCUMULATESOnly one field has a reducer, and it is the one you want to grow. The default behaviour for the others is overwrite, which is exactly right for them.
Get this wrong and the symptom is silent. Suppose trace had no reducer and four nodes each returned {"trace": ["classified as math"]}, {"trace": ["computed 61.2"]} and so on. The final state holds one entry, not four, and there is no error to tell you.
| Field | Written by | Wanted behaviour | Annotation |
|---|---|---|---|
question | The caller | Set once | none |
route | classify | Latest wins | none |
answer | Three branch nodes | Latest wins — a retry replaces the old answer | none |
attempts | Branch nodes | Latest wins (node computes old + 1) | none |
trace | Every node | Grow | operator.add |
Before writing a single node, go through the state schema field by field and answer one question for each: if two nodes write this, should the second replace the first or add to it? That decision is the reducer, and it is unrecoverable once the run has finished.
Requirement 2 — the classifier node
A node is a plain function. Take state, return a dict of updates. Nothing more.
1VALID = {"math", "search", "code"}23def classify(state: State) -> dict:4 raw = fast.invoke(5 "Classify the user's question into exactly one category.\n"6 "math - arithmetic, percentages, unit conversion, numeric comparison\n"7 "search - facts that may have changed: news, prices, people's roles\n"8 "code - reading, explaining, debugging or writing source code\n"9 "Reply with the single word and nothing else.\n\n"10 f"Question: {state['question']}"11 ).content.strip().lower()1213 route = raw if raw in VALID else "search"14 return {"route": route,15 "attempts": 0,16 "trace": [f"classify: raw={raw!r} -> {route}"]}The line route = raw if raw in VALID else "search" is the important one. Models occasionally reply "Math." or "This is a math question", and an unrecognised label passed into a conditional edge raises a KeyError that kills the run. Validating against a set and falling back to the safest branch converts a crash into a slightly suboptimal route.
Choose your fallback deliberately. search is the right default here because it is the most general branch — it can produce a reasonable answer to a math or code question, whereas the calculator branch cannot answer "who is the CFO of Siemens".
Requirement 3 — the three branches
1_OPS = {ast.Add: operator.add, ast.Sub: operator.sub, ast.Mult: operator.mul,2 ast.Div: operator.truediv, ast.Pow: operator.pow, ast.USub: operator.neg}34def _calc(node):5 if isinstance(node, ast.Constant) and isinstance(node.value, (int, float)):6 return node.value7 if isinstance(node, ast.BinOp) and type(node.op) in _OPS:8 return _OPS[type(node.op)](_calc(node.left), _calc(node.right))9 if isinstance(node, ast.UnaryOp) and type(node.op) in _OPS:10 return _OPS[type(node.op)](_calc(node.operand))11 raise ValueError("unsupported expression")1213def math_node(state: State) -> dict:14 expr = fast.invoke(15 "Rewrite this as a single arithmetic expression using only digits and "16 "+ - * / ( ) ** . No words, no units, no percent signs. "17 "Write percentages as decimals.\n\n" + state["question"]18 ).content.strip().strip("`")19 try:20 value = _calc(ast.parse(expr, mode="eval").body)21 answer = f"{state['question'].rstrip('?')} = {value:,.4f}".rstrip("0").rstrip(".")22 except Exception as e:23 answer = f"Could not evaluate {expr!r}: {e}"24 return {"answer": answer,25 "attempts": state["attempts"] + 1,26 "trace": [f"math: expr={expr!r}"]}2728def search_node(state: State) -> dict:29 from ddgs import DDGS30 with DDGS() as d:31 hits = list(d.text(state["question"], max_results=4))32 context = "\n".join(f"- {h['title']}: {h['body'][:200]}" for h in hits) or "none"33 answer = smart.invoke(34 "Answer the question using ONLY these snippets. If they do not contain "35 "the answer, say so plainly.\n\n"36 f"Snippets:\n{context}\n\nQuestion: {state['question']}"37 ).content38 return {"answer": answer,39 "attempts": state["attempts"] + 1,40 "trace": [f"search: {len(hits)} hits"]}4142def code_node(state: State) -> dict:43 answer = smart.invoke(44 "You are a senior engineer. Answer precisely. If the question contains "45 "code, quote the exact line that causes the behaviour before explaining "46 "it.\n\n" + state["question"]47 ).content48 return {"answer": answer,49 "attempts": state["attempts"] + 1,50 "trace": ["code: answered"]}Each branch does one thing and returns an answer. None of them decide what runs next — that is the graph's job, and keeping it out of the node functions is what makes them individually testable.
Requirement 4 — wiring the graph
Add a validation node and a cycle, so the workflow contains a genuine loop with a hard limit rather than an instruction that hopes to be obeyed.
1def validate(state: State) -> dict:2 v = fast.invoke(3 "Does this answer actually address the question? "4 "Reply PASS, or one short sentence naming what is missing.\n\n"5 f"Question: {state['question']}\nAnswer: {state['answer']}"6 ).content.strip()7 return {"verdict": v, "trace": [f"validate: {v[:60]}"]}89def to_branch(state: State) -> Literal["math", "search", "code"]:10 return state["route"]1112def after_validate(state: State) -> Literal["done", "retry"]:13 if state["verdict"].upper().startswith("PASS"):14 return "done"15 if state["attempts"] >= 2:16 return "done" # give up gracefully rather than loop17 return "retry"1819g = StateGraph(State)20g.add_node("classify", classify)21g.add_node("math", math_node)22g.add_node("search", search_node)23g.add_node("code", code_node)24g.add_node("validate", validate)2526g.add_edge(START, "classify")27g.add_conditional_edges("classify", to_branch,28 {"math": "math", "search": "search", "code": "code"})29for b in ("math", "search", "code"):30 g.add_edge(b, "validate")31g.add_conditional_edges("validate", after_validate,32 {"done": END, "retry": "search"})3334app = g.compile()The retry edge points at search rather than back at the original branch on purpose: if the first attempt failed validation, repeating the same branch usually reproduces the same failure, whereas a web search brings in information the first attempt did not have. That is a design choice, not a rule — send it back to to_branch instead if you prefer, but then the counter is doing all the work preventing an infinite loop.
The shape of the compiled graph:
START | classify / | \ math search code \ | / validate / \ PASS / \ FAIL and attempts < 2 | | END search (loop back, max once)A conditional edge is only as safe as its fallback. Any label the mapping does not contain is a crash, so validate the classifier's answer before the graph ever sees it.
Testing it
1CASES = [2 ("What is 18% of 340?", "math"),3 ("What is 847 * 293?", "math"),4 ("Who is the current CEO of Siemens?", "search"),5 ("What happened to oil prices this week?", "search"),6 ("Why does this return None?\n\ndef f(x):\n x.sort()\n return x.sort()",7 "code"),8 ("Explain Python's GIL in two sentences.", "code"),9]1011hits = 012for q, expected in CASES:13 out = app.invoke({"question": q})14 ok = out["route"] == expected15 hits += ok16 print(f"{'OK ' if ok else 'MISS'} route={out['route']:<7} expected={expected}")17 for line in out["trace"]:18 print(" " + line)19print(f"routing accuracy: {hits}/{len(CASES)}")Two of these have answers you can check by hand, which is the point of including them. 0.18×340=61.2 and 847×293=248,171. If the math branch returns anything else, the expression-rewriting prompt is producing something the parser mangles — print expr from the trace and you will see it immediately, usually a stray % or a comma inside a number.
The third-from-last case is the interesting one. x.sort() sorts in place and returns None, so return x.sort() returns None — the code branch should quote that line. If it gets routed to search instead, your classifier prompt needs the word "debugging" made more prominent, because the model is latching onto "why does" as a question about the world.
Where this goes wrong
| Symptom | Cause | Fix |
|---|---|---|
KeyError: 'Math.' from the conditional edge | Classifier returned a label not in the mapping | Validate against a set, fall back to a default |
trace has one entry at the end | Missing Annotated[list, operator.add] | Add the reducer |
GraphRecursionError | Retry edge with no attempt counter, or attempts reset inside the branch | Increment in the branch, check in after_validate |
| Node's change is invisible downstream | Node mutated state["x"] = ... instead of returning | Always return {...} |
Everything routes to search | Classifier replying in a sentence, hitting the fallback every time | Print the raw reply; add "and nothing else" to the prompt |
| Retry never triggers | Validator almost always says PASS | Make the validator's criterion specific and checkable, not "is this good" |
Worth extending
- Add a fourth branch and measure the classifier again. Routing accuracy falls as categories multiply and their boundaries blur. Watching that happen with your own numbers is more convincing than being told.
- Fan out instead of routing. For ambiguous questions, run
searchandcodein parallel by adding both edges out ofclassify, then add a merge node. You will need a reducer onanswertoo — a list rather than a string. - Compile with a checkpointer (
SqliteSaver, fromlanggraph-checkpoint-sqlite) and athread_id. Kill the process mid-run and resume it. The state comes back. - Stream it:
for ev in app.stream(inp, stream_mode="updates")gives you per-node progress, which is what you would show a user as "Classifying… Searching… Checking answer…".
What the graph bought you
Compare the finished system with the single agent described at the start. The agent carried all three tools in every request; the graph carries one branch's worth. On a math question that is roughly 180 tokens of schema instead of 540 — small per request, meaningful at 50,000 requests a month.
But token savings are the lesser prize. The real difference is that three properties of this workflow are now facts about the code rather than hopes about the model: a question is handled by exactly one branch, every answer passes through validation, and no request costs more than two attempts. You can read those guarantees off the graph definition and prove them with a unit test on after_validate that never calls a model at all.
That is the trade a graph asks you to make. You give up the model's freedom to improvise a sequence of steps, and you buy a workflow whose control flow you can point at, test and change. Take that trade when the steps are known and the guarantees matter. Refuse it when the whole value of the system is that you did not know in advance what it would need to do.