Course Content
AI Agent Frameworks
4 sections · 15 lessons
LangGraph and Graph-based Workflows
A content team asked for something that sounds trivial: an agent that drafts a product description, scores it against a rubric, and rewrites it if the score is below 7 out of 10 — up to three attempts, and if it still fails, hand it to a human.
Try to express that with a standard tool-calling agent loop. You can give the model a draft tool and a score tool and hope it works out the protocol from the system prompt. What you get in practice, from a real run:
Attempt 1: draft -> score 5 -> redraftAttempt 2: draft -> score 6 -> redraftAttempt 3: draft -> score 6 -> redraftAttempt 4: draft -> score 7 -> "Here is your description!"Four attempts, not three. And on a different run the model scored its own draft, decided 6 was "close enough", and returned it. On a third run it forgot to score at all. The retry limit lived in an English sentence in the prompt, so it was a suggestion rather than a rule.
The problem is structural. A tool-calling agent is a loop with one decision point: call a tool, or stop. Everything else — the order of steps, the retry count, the branch conditions — is smuggled into prose and enforced by hope. When your workflow genuinely has branches, cycles with hard limits, or points where a human must approve something, prose is the wrong place to put the control flow.
The moment your workflow contains the words "if", "until", or "at most N times", the control flow belongs in code, not in a prompt.
LangGraph is the answer to exactly that. You define your workflow as a directed graph: nodes are functions, edges are transitions, and some edges are conditional. The model decides content; the graph decides flow.
The three primitives
pip install langgraph langchain-anthropicEverything in LangGraph is built from three things.
| Primitive | What it is | In code |
|---|---|---|
| State | A typed dictionary passed to every node and updated by each one | A TypedDict |
| Node | A plain function: takes state, returns a partial state update | def n(state) -> dict |
| Edge | Which node runs next — fixed, or chosen by a function | add_edge / add_conditional_edges |
Here is the draft-score-revise workflow, with the retry limit as an actual limit.
1from typing import TypedDict, Literal2from langgraph.graph import StateGraph, START, END34class State(TypedDict):5 brief: str6 draft: str7 score: int8 feedback: str9 attempts: int1011def draft_node(state: State) -> dict:12 if state.get("feedback"):13 prompt = (f"Rewrite this product description.\n"14 f"Brief: {state['brief']}\n"15 f"Previous draft: {state['draft']}\n"16 f"Reviewer feedback: {state['feedback']}")17 else:18 prompt = f"Write a product description.\nBrief: {state['brief']}"19 return {"draft": llm.invoke(prompt).content,20 "attempts": state.get("attempts", 0) + 1}2122def score_node(state: State) -> dict:23 out = llm.invoke(24 "Score this description 1-10 for clarity, specificity and tone. "25 "Reply as: SCORE: n | FEEDBACK: one sentence\n\n" + state["draft"]26 ).content27 score = int(out.split("SCORE:")[1].split("|")[0].strip())28 feedback = out.split("FEEDBACK:")[1].strip()29 return {"score": score, "feedback": feedback}3031def route(state: State) -> Literal["accept", "revise", "escalate"]:32 if state["score"] >= 7:33 return "accept"34 if state["attempts"] >= 3:35 return "escalate"36 return "revise"3738def accept_node(state: State) -> dict:39 return {"feedback": "APPROVED"}4041def escalate_node(state: State) -> dict:42 return {"feedback": f"HUMAN REVIEW NEEDED after {state['attempts']} attempts"}4344g = StateGraph(State)45g.add_node("draft", draft_node)46g.add_node("score", score_node)47g.add_node("accept", accept_node)48g.add_node("escalate", escalate_node)4950g.add_edge(START, "draft")51g.add_edge("draft", "score")52g.add_conditional_edges("score", route,53 {"accept": "accept", "revise": "draft",54 "escalate": "escalate"})55g.add_edge("accept", END)56g.add_edge("escalate", END)5758app = g.compile()59result = app.invoke({"brief": "Waterproof hiking boot, 240 g, vegan upper",60 "attempts": 0})Now "at most three attempts" is state["attempts"] >= 3. It is enforced by Python, not by the model's willingness to count. The model still writes and judges the copy — the parts that need judgement — but it has no say in whether attempt four happens.
Note also that route returns a label, not a node name, and the mapping dictionary translates labels to nodes. That indirection lets you rewire the graph without touching the routing logic, and it makes the routing function trivially unit-testable: feed it a state dict, assert the label.
State: the part people get wrong
A node returns a partial update. LangGraph merges it into the running state. The default merge is overwrite — the returned value replaces the old one. That is correct for draft and score, and catastrophically wrong for anything you want to accumulate.
class Bad(TypedDict): findings: list[str] # each node returns {"findings": [...]} # -> later nodes ERASE earlier findingsTo accumulate, annotate the field with a reducer — a function of (old, new) that says how to combine them.
1from typing import Annotated2import operator3from langgraph.graph.message import add_messages45class Good(TypedDict):6 question: str # overwrite (default)7 findings: Annotated[list[str], operator.add] # concatenate8 cost: Annotated[float, operator.add] # sum9 messages: Annotated[list, add_messages] # append, dedupe by idWalk the arithmetic once so the behaviour is unambiguous. Three research nodes run and each returns {"findings": ["..."], "cost": 0.004}.
| Field | Reducer | After node 1 | After node 2 | After node 3 |
|---|---|---|---|---|
findings | operator.add | 1 item | 2 items | 3 items |
cost | operator.add | 0.004 | 0.008 | 0.012 |
findings | none (default) | 1 item | 1 item | 1 item |
The bottom row is the bug: you run three researchers, pay for three, and keep one. It produces no error, no warning — just a synthesis step working from a third of the evidence. This is the single most common LangGraph mistake, and it is invisible unless you check the state.
Every field in your state schema needs a deliberate answer to one question: when two nodes write to this, should the second replace the first or be added to it?
Routing patterns worth knowing
Loop until converged
The draft-score-revise cycle above. The essential ingredient is a counter in state and a check against it — a cycle with no counter is an infinite loop, and LangGraph will happily run it until the recursion limit trips. In current LangGraph that default limit is in the thousands of steps, so set your own.
app = g.compile()app.invoke(initial, config={"recursion_limit": 25}) # safety net, not a designParallel branches (fan-out and fan-in)
Add several edges out of one node and LangGraph runs those nodes concurrently, then waits for all of them before running any node downstream of all of them.
1g.add_edge("plan", "search_news")2g.add_edge("plan", "search_academic")3g.add_edge("plan", "search_internal")45g.add_edge("search_news", "synthesise")6g.add_edge("search_academic", "synthesise")7g.add_edge("search_internal", "synthesise")The timing difference is real. If each search takes 4 seconds and synthesis takes 3, sequential execution is 3×4+3=15 seconds; parallel is max(4,4,4)+3=7 seconds. Fan-in is where the reducer matters most: all three searches write to findings, so without operator.add you keep whichever finished last.
Dynamic fan-out with Send
When you do not know how many branches you need until runtime — one per sub-question the planner generated — use Send.
1from langgraph.types import Send23def dispatch(state: State):4 return [Send("research_one", {"subq": q}) for q in state["subquestions"]]56g.add_conditional_edges("plan", dispatch, ["research_one"])7g.add_edge("research_one", "synthesise")Five sub-questions produce five parallel invocations of research_one, each with its own slice of input, all merging into shared state through the reducer.
Classify then dispatch
A cheap first node labels the request; a conditional edge sends it down one of several specialised paths. This is how you avoid loading twenty tools into every request.
1def classify(state) -> dict:2 label = small_llm.invoke(3 "One word - math, search or code:\n" + state["question"]).content.strip()4 return {"route": label if label in {"math", "search", "code"} else "search"}56g.add_conditional_edges("classify", lambda s: s["route"],7 {"math": "math_node", "search": "search_node",8 "code": "code_node"})Human in the loop
Some steps must not happen without a person's approval: sending an email, issuing a refund, running a migration. LangGraph handles this by interrupting — the graph stops before a named node, persists its state, and returns. Later, possibly in a different process, you resume.
1from langgraph.checkpoint.memory import InMemorySaver23app = g.compile(checkpointer=InMemorySaver(), interrupt_before=["send_email"])45cfg = {"configurable": {"thread_id": "ticket-9931"}}6state = app.invoke({"request": "Refund order 4471"}, cfg)78snapshot = app.get_state(cfg)9print(snapshot.next) # ('send_email',)10print(snapshot.values["draft"]) # show the human what is about to be sent1112# a human edits and approves...13app.update_state(cfg, {"draft": edited_text})14final = app.invoke(None, cfg) # None = resume from the checkpointTwo things make this work. The checkpointer saves state after every node, so the run can survive a process restart. The thread_id identifies the conversation, so resuming means "continue thread ticket-9931", not "re-run everything". InMemorySaver is for development; SqliteSaver or PostgresSaver is what you deploy.
interrupt_before is the simplest way to see the mechanism. For approval steps in a real product, current LangGraph documentation recommends calling interrupt() inside the node instead: the node pauses at that line, hands the draft to the caller, and continues with whatever Command(resume=...) value the human sends back. The checkpointer and thread_id work the same way in both styles.
The same machinery gives you three other capabilities for free: crash recovery (resume from the last completed node), time travel (rewind to an earlier checkpoint and take a different branch), and full conversation persistence across days.
Streaming what is happening
for event in app.stream(initial, cfg, stream_mode="updates"): for node, update in event.items(): print(f"{node}: {list(update.keys())}")stream_mode | Yields | Use for |
|---|---|---|
"updates" | Each node's partial update | Progress indicators: "Searching news…" |
"values" | The full state after each node | Debugging state evolution |
"messages" | Token-by-token model output | Typing-effect UIs |
Subgraphs
A compiled graph is itself a valid node. That is how you keep a large workflow readable: build and test a research sub-workflow on its own, then drop it into the parent.
1research_app = research_graph.compile()23parent = StateGraph(ParentState)4parent.add_node("research", research_app) # a whole graph as one node5parent.add_node("write", write_node)6parent.add_edge(START, "research")7parent.add_edge("research", "write")The subgraph must be able to read and write the state keys the parent passes it. If their schemas differ, wrap the subgraph in a function that translates keys in and out — that translation layer is much easier to debug than a silently missing field.
A complete research workflow
1# pip install langgraph-checkpoint-sqlite2from typing import TypedDict, Annotated3import operator, sqlite34from langgraph.graph import StateGraph, START, END5from langgraph.types import Send6from langgraph.checkpoint.sqlite import SqliteSaver78class RState(TypedDict):9 question: str10 subquestions: list[str]11 findings: Annotated[list[str], operator.add]12 report: str13 critique: str14 revisions: int1516def plan(s):17 qs = llm.invoke(f"List 3 sub-questions, one per line:\n{s['question']}")18 return {"subquestions": [l.strip("- ") for l in qs.content.splitlines()19 if l.strip()][:3], "revisions": 0}2021def research_one(s):22 hits = search(s["subq"])23 return {"findings": [f"Q: {s['subq']}\nA: {hits}"]}2425def synthesise(s):26 body = "\n\n".join(s["findings"])27 return {"report": llm.invoke(28 f"Write a 300-word report on {s['question']} using ONLY:\n{body}").content}2930def critique(s):31 c = llm.invoke("Reply PASS, or one specific problem:\n" + s["report"]).content32 return {"critique": c, "revisions": s["revisions"] + 1}3334def after_critique(s):35 if s["critique"].startswith("PASS") or s["revisions"] >= 2:36 return "done"37 return "redo"3839g = StateGraph(RState)40for name, fn in [("plan", plan), ("research_one", research_one),41 ("synthesise", synthesise), ("critique", critique)]:42 g.add_node(name, fn)4344g.add_edge(START, "plan")45g.add_conditional_edges("plan",46 lambda s: [Send("research_one", {"subq": q, "question": s["question"]})47 for q in s["subquestions"]], ["research_one"])48g.add_edge("research_one", "synthesise")49g.add_edge("synthesise", "critique")50g.add_conditional_edges("critique", after_critique,51 {"done": END, "redo": "synthesise"})5253app = g.compile(checkpointer=SqliteSaver(54 sqlite3.connect("runs.db", check_same_thread=False)))Read the shape: plan once, fan out to three parallel researchers, fan in to a synthesiser, critique, and loop back to synthesis at most once (the second critique ends the run whatever it says). Every one of those constraints is a line of Python. None of them depend on the model remembering an instruction.
When a graph is the wrong tool
| Tool-calling agent loop | Graph workflow | |
|---|---|---|
| Who chooses the order of steps | The model, per request | You, at design time |
| Cycles with hard limits | Prompt instruction (advisory) | Counter in state (enforced) |
| Parallel work | Only parallel tool calls in one step | Arbitrary parallel branches |
| Pause and resume days later | No | Yes, via checkpointer |
| Lines of setup for one tool | ~10 | ~35 |
| Debugging | Read the trace, infer intent | Inspect state at each node |
| Right when | Steps are unpredictable and few | Steps are known and structured |
The honest summary: if your task is "answer a question, using whichever of these four tools seems relevant", a graph is overhead you do not need. The model's flexibility is the feature. If your task is "do A, then either B or C, then D, retrying D up to twice, and pause for approval before E", a plain agent loop will approximate that and get it wrong perhaps one run in five.
Named failure modes
| Symptom | Cause | Fix |
|---|---|---|
| Parallel branches produce one result | No reducer on the accumulating field | Annotated[list, operator.add] |
GraphRecursionError | Cycle with no counter, or a routing function that never returns the exit label | Add an attempt counter and check it first |
| Node sees stale values | Node mutated the state dict in place instead of returning an update | Always return {...}; never assign to state[...] |
| Resume restarts from the beginning | Missing or changed thread_id, or no checkpointer | Stable thread ID, persistent checkpointer |
| Graph is unreadable at 20 nodes | Everything in one flat graph | Extract cohesive regions into subgraphs |
Designing the graph before you write it
The productive habit is to draw the graph on paper first, and for each node write down two things: what it reads from state, and what it writes. That two-column list is your state schema, and going through it forces the reducer question on every field before it can bite you.
Then ask, for each edge, whether it is fixed or conditional — and for each conditional edge, what the exit condition is when things go badly. A cycle without an answer to "what if this never converges?" is a production incident waiting for its trigger.
Finally, keep nodes narrow. A node that calls a model, parses the result, writes to a database and decides where to go next is four responsibilities in one function, and when the run misbehaves you cannot tell which one failed. Separate node functions that each do one thing let you inspect state between them, unit-test the routing logic without a model, and replace one step without touching the rest of the graph. That granularity is the actual payoff of expressing a workflow as a graph — not the picture, but the fact that every transition is a thing you can point at, test and change.