AI Agent Frameworks

LangGraph and Graph-based Workflows


A content team asked for something that sounds trivial: an agent that drafts a product description, scores it against a rubric, and rewrites it if the score is below 7 out of 10 — up to three attempts, and if it still fails, hand it to a human.

Try to express that with a standard tool-calling agent loop. You can give the model a draft tool and a score tool and hope it works out the protocol from the system prompt. What you get in practice, from a real run:

Text
Attempt 1: draft -> score 5  -> redraftAttempt 2: draft -> score 6  -> redraftAttempt 3: draft -> score 6  -> redraftAttempt 4: draft -> score 7  -> "Here is your description!"

Four attempts, not three. And on a different run the model scored its own draft, decided 6 was "close enough", and returned it. On a third run it forgot to score at all. The retry limit lived in an English sentence in the prompt, so it was a suggestion rather than a rule.

The problem is structural. A tool-calling agent is a loop with one decision point: call a tool, or stop. Everything else — the order of steps, the retry count, the branch conditions — is smuggled into prose and enforced by hope. When your workflow genuinely has branches, cycles with hard limits, or points where a human must approve something, prose is the wrong place to put the control flow.

The moment your workflow contains the words "if", "until", or "at most N times", the control flow belongs in code, not in a prompt.

LangGraph is the answer to exactly that. You define your workflow as a directed graph: nodes are functions, edges are transitions, and some edges are conditional. The model decides content; the graph decides flow.

Draft, score, and the loop backdraftscorepublishscore under 7redraft if under 3hand to human at 3
The cycle is legal because the attempt count lives in state — without it the same graph loops forever.

The three primitives

Bash
pip install langgraph langchain-anthropic

Everything in LangGraph is built from three things.

PrimitiveWhat it isIn code
StateA typed dictionary passed to every node and updated by each oneA TypedDict
NodeA plain function: takes state, returns a partial state updatedef n(state) -> dict
EdgeWhich node runs next — fixed, or chosen by a functionadd_edge / add_conditional_edges

Here is the draft-score-revise workflow, with the retry limit as an actual limit.

Python
from typing import TypedDict, Literalfrom langgraph.graph import StateGraph, START, ENDclass State(TypedDict):    brief: str    draft: str    score: int    feedback: str    attempts: intdef draft_node(state: State) -> dict:    if state.get("feedback"):        prompt = (f"Rewrite this product description.\n"                  f"Brief: {state['brief']}\n"                  f"Previous draft: {state['draft']}\n"                  f"Reviewer feedback: {state['feedback']}")    else:        prompt = f"Write a product description.\nBrief: {state['brief']}"    return {"draft": llm.invoke(prompt).content,            "attempts": state.get("attempts", 0) + 1}def score_node(state: State) -> dict:    out = llm.invoke(        "Score this description 1-10 for clarity, specificity and tone. "        "Reply as: SCORE: n | FEEDBACK: one sentence\n\n" + state["draft"]    ).content    score = int(out.split("SCORE:")[1].split("|")[0].strip())    feedback = out.split("FEEDBACK:")[1].strip()    return {"score": score, "feedback": feedback}def route(state: State) -> Literal["accept", "revise", "escalate"]:    if state["score"] >= 7:        return "accept"    if state["attempts"] >= 3:        return "escalate"    return "revise"def accept_node(state: State) -> dict:    return {"feedback": "APPROVED"}def escalate_node(state: State) -> dict:    return {"feedback": f"HUMAN REVIEW NEEDED after {state['attempts']} attempts"}g = StateGraph(State)g.add_node("draft", draft_node)g.add_node("score", score_node)g.add_node("accept", accept_node)g.add_node("escalate", escalate_node)g.add_edge(START, "draft")g.add_edge("draft", "score")g.add_conditional_edges("score", route,                        {"accept": "accept", "revise": "draft",                         "escalate": "escalate"})g.add_edge("accept", END)g.add_edge("escalate", END)app = g.compile()result = app.invoke({"brief": "Waterproof hiking boot, 240 g, vegan upper",                     "attempts": 0})

Now "at most three attempts" is state["attempts"] >= 3. It is enforced by Python, not by the model's willingness to count. The model still writes and judges the copy — the parts that need judgement — but it has no say in whether attempt four happens.

Note also that route returns a label, not a node name, and the mapping dictionary translates labels to nodes. That indirection lets you rewire the graph without touching the routing logic, and it makes the routing function trivially unit-testable: feed it a state dict, assert the label.

State: the part people get wrong

A node returns a partial update. LangGraph merges it into the running state. The default merge is overwrite — the returned value replaces the old one. That is correct for draft and score, and catastrophically wrong for anything you want to accumulate.

Python
class Bad(TypedDict):    findings: list[str]     # each node returns {"findings": [...]}                            # -> later nodes ERASE earlier findings

To accumulate, annotate the field with a reducer — a function of (old, new) that says how to combine them.

Python
from typing import Annotatedimport operatorfrom langgraph.graph.message import add_messagesclass Good(TypedDict):    question: str                                   # overwrite (default)    findings: Annotated[list[str], operator.add]    # concatenate    cost: Annotated[float, operator.add]            # sum    messages: Annotated[list, add_messages]         # append, dedupe by id

Walk the arithmetic once so the behaviour is unambiguous. Three research nodes run and each returns {"findings": ["..."], "cost": 0.004}.

FieldReducerAfter node 1After node 2After node 3
findingsoperator.add1 item2 items3 items
costoperator.add0.0040.0080.012
findingsnone (default)1 item1 item1 item

The bottom row is the bug: you run three researchers, pay for three, and keep one. It produces no error, no warning — just a synthesis step working from a third of the evidence. This is the single most common LangGraph mistake, and it is invisible unless you check the state.

Every field in your state schema needs a deliberate answer to one question: when two nodes write to this, should the second replace the first or be added to it?

Routing patterns worth knowing

Loop until converged

The draft-score-revise cycle above. The essential ingredient is a counter in state and a check against it — a cycle with no counter is an infinite loop, and LangGraph will happily run it until the recursion limit trips. In current LangGraph that default limit is in the thousands of steps, so set your own.

Python
app = g.compile()app.invoke(initial, config={"recursion_limit": 25})   # safety net, not a design

Parallel branches (fan-out and fan-in)

Add several edges out of one node and LangGraph runs those nodes concurrently, then waits for all of them before running any node downstream of all of them.

Python
g.add_edge("plan", "search_news")g.add_edge("plan", "search_academic")g.add_edge("plan", "search_internal")g.add_edge("search_news", "synthesise")g.add_edge("search_academic", "synthesise")g.add_edge("search_internal", "synthesise")

The timing difference is real. If each search takes 4 seconds and synthesis takes 3, sequential execution is 3×4+3=153 \times 4 + 3 = 15 seconds; parallel is max⁡(4,4,4)+3=7\max(4,4,4) + 3 = 7 seconds. Fan-in is where the reducer matters most: all three searches write to findings, so without operator.add you keep whichever finished last.

Dynamic fan-out with Send

When you do not know how many branches you need until runtime — one per sub-question the planner generated — use Send.

Python
from langgraph.types import Senddef dispatch(state: State):    return [Send("research_one", {"subq": q}) for q in state["subquestions"]]g.add_conditional_edges("plan", dispatch, ["research_one"])g.add_edge("research_one", "synthesise")

Five sub-questions produce five parallel invocations of research_one, each with its own slice of input, all merging into shared state through the reducer.

Classify then dispatch

A cheap first node labels the request; a conditional edge sends it down one of several specialised paths. This is how you avoid loading twenty tools into every request.

Python
def classify(state) -> dict:    label = small_llm.invoke(        "One word - math, search or code:\n" + state["question"]).content.strip()    return {"route": label if label in {"math", "search", "code"} else "search"}g.add_conditional_edges("classify", lambda s: s["route"],                        {"math": "math_node", "search": "search_node",                         "code": "code_node"})

Human in the loop

Some steps must not happen without a person's approval: sending an email, issuing a refund, running a migration. LangGraph handles this by interrupting — the graph stops before a named node, persists its state, and returns. Later, possibly in a different process, you resume.

Python
from langgraph.checkpoint.memory import InMemorySaverapp = g.compile(checkpointer=InMemorySaver(), interrupt_before=["send_email"])cfg = {"configurable": {"thread_id": "ticket-9931"}}state = app.invoke({"request": "Refund order 4471"}, cfg)snapshot = app.get_state(cfg)print(snapshot.next)            # ('send_email',)print(snapshot.values["draft"]) # show the human what is about to be sent# a human edits and approves...app.update_state(cfg, {"draft": edited_text})final = app.invoke(None, cfg)   # None = resume from the checkpoint

Two things make this work. The checkpointer saves state after every node, so the run can survive a process restart. The thread_id identifies the conversation, so resuming means "continue thread ticket-9931", not "re-run everything". InMemorySaver is for development; SqliteSaver or PostgresSaver is what you deploy.

interrupt_before is the simplest way to see the mechanism. For approval steps in a real product, current LangGraph documentation recommends calling interrupt() inside the node instead: the node pauses at that line, hands the draft to the caller, and continues with whatever Command(resume=...) value the human sends back. The checkpointer and thread_id work the same way in both styles.

The same machinery gives you three other capabilities for free: crash recovery (resume from the last completed node), time travel (rewind to an earlier checkpoint and take a different branch), and full conversation persistence across days.

Streaming what is happening

Python
for event in app.stream(initial, cfg, stream_mode="updates"):    for node, update in event.items():        print(f"{node}: {list(update.keys())}")
stream_modeYieldsUse for
"updates"Each node's partial updateProgress indicators: "Searching news…"
"values"The full state after each nodeDebugging state evolution
"messages"Token-by-token model outputTyping-effect UIs

Subgraphs

A compiled graph is itself a valid node. That is how you keep a large workflow readable: build and test a research sub-workflow on its own, then drop it into the parent.

Python
research_app = research_graph.compile()parent = StateGraph(ParentState)parent.add_node("research", research_app)     # a whole graph as one nodeparent.add_node("write", write_node)parent.add_edge(START, "research")parent.add_edge("research", "write")

The subgraph must be able to read and write the state keys the parent passes it. If their schemas differ, wrap the subgraph in a function that translates keys in and out — that translation layer is much easier to debug than a silently missing field.

A complete research workflow

Python
# pip install langgraph-checkpoint-sqlitefrom typing import TypedDict, Annotatedimport operator, sqlite3from langgraph.graph import StateGraph, START, ENDfrom langgraph.types import Sendfrom langgraph.checkpoint.sqlite import SqliteSaverclass RState(TypedDict):    question: str    subquestions: list[str]    findings: Annotated[list[str], operator.add]    report: str    critique: str    revisions: intdef plan(s):    qs = llm.invoke(f"List 3 sub-questions, one per line:\n{s['question']}")    return {"subquestions": [l.strip("- ") for l in qs.content.splitlines()                             if l.strip()][:3], "revisions": 0}def research_one(s):    hits = search(s["subq"])    return {"findings": [f"Q: {s['subq']}\nA: {hits}"]}def synthesise(s):    body = "\n\n".join(s["findings"])    return {"report": llm.invoke(        f"Write a 300-word report on {s['question']} using ONLY:\n{body}").content}def critique(s):    c = llm.invoke("Reply PASS, or one specific problem:\n" + s["report"]).content    return {"critique": c, "revisions": s["revisions"] + 1}def after_critique(s):    if s["critique"].startswith("PASS") or s["revisions"] >= 2:        return "done"    return "redo"g = StateGraph(RState)for name, fn in [("plan", plan), ("research_one", research_one),                 ("synthesise", synthesise), ("critique", critique)]:    g.add_node(name, fn)g.add_edge(START, "plan")g.add_conditional_edges("plan",    lambda s: [Send("research_one", {"subq": q, "question": s["question"]})               for q in s["subquestions"]], ["research_one"])g.add_edge("research_one", "synthesise")g.add_edge("synthesise", "critique")g.add_conditional_edges("critique", after_critique,                        {"done": END, "redo": "synthesise"})app = g.compile(checkpointer=SqliteSaver(    sqlite3.connect("runs.db", check_same_thread=False)))

Read the shape: plan once, fan out to three parallel researchers, fan in to a synthesiser, critique, and loop back to synthesis at most once (the second critique ends the run whatever it says). Every one of those constraints is a line of Python. None of them depend on the model remembering an instruction.

When a graph is the wrong tool

Tool-calling agent loopGraph workflow
Who chooses the order of stepsThe model, per requestYou, at design time
Cycles with hard limitsPrompt instruction (advisory)Counter in state (enforced)
Parallel workOnly parallel tool calls in one stepArbitrary parallel branches
Pause and resume days laterNoYes, via checkpointer
Lines of setup for one tool~10~35
DebuggingRead the trace, infer intentInspect state at each node
Right whenSteps are unpredictable and fewSteps are known and structured

The honest summary: if your task is "answer a question, using whichever of these four tools seems relevant", a graph is overhead you do not need. The model's flexibility is the feature. If your task is "do A, then either B or C, then D, retrying D up to twice, and pause for approval before E", a plain agent loop will approximate that and get it wrong perhaps one run in five.

Named failure modes

SymptomCauseFix
Parallel branches produce one resultNo reducer on the accumulating fieldAnnotated[list, operator.add]
GraphRecursionErrorCycle with no counter, or a routing function that never returns the exit labelAdd an attempt counter and check it first
Node sees stale valuesNode mutated the state dict in place instead of returning an updateAlways return {...}; never assign to state[...]
Resume restarts from the beginningMissing or changed thread_id, or no checkpointerStable thread ID, persistent checkpointer
Graph is unreadable at 20 nodesEverything in one flat graphExtract cohesive regions into subgraphs

Designing the graph before you write it

The productive habit is to draw the graph on paper first, and for each node write down two things: what it reads from state, and what it writes. That two-column list is your state schema, and going through it forces the reducer question on every field before it can bite you.

Then ask, for each edge, whether it is fixed or conditional — and for each conditional edge, what the exit condition is when things go badly. A cycle without an answer to "what if this never converges?" is a production incident waiting for its trigger.

Finally, keep nodes narrow. A node that calls a model, parses the result, writes to a database and decides where to go next is four responsibilities in one function, and when the run misbehaves you cannot tell which one failed. Separate node functions that each do one thing let you inspect state between them, unit-test the routing logic without a model, and replace one step without touching the rest of the graph. That granularity is the actual payoff of expressing a workflow as a graph — not the picture, but the fact that every transition is a thing you can point at, test and change.