Multi-Agent Systems and Collaboration

LangGraph for Multi-Agent Flows


A team hand-rolled a three-agent content pipeline: a researcher, a writer, and a critic. The orchestrator was a single Python function. It started at 40 lines. Then the product owner asked for one more behaviour — if the critic rejects the draft, send it back to the writer, up to three times.

That function grew to 220 lines and six boolean flags: has_research, has_draft, needs_revision, critic_ran, escalated, done. Then a bug appeared. needs_revision was set inside the retry branch but never cleared when the writer succeeded, so the loop ran until the step cap: 47 model calls for one blog post, roughly forty of them redundant.

The bug was trivial to fix and the shape of the code guaranteed it would come back. Control flow lived in flags, state lived in local variables, and there was no way to see what had happened except by reading logs. When the process was killed at call 30, everything was lost — there was nothing to resume from, because the state existed only in a stack frame.

That is the problem a graph orchestration layer solves, and it is worth being precise about what it actually gives you, because "you can draw it as a diagram" is not the reason.

Researcher, writer, critic — and the rejection edgeresearcherwritercriticrevisioncountpublishrejectsthe draftcapped at threeState is the message bus; the reducer decides whether a field is replaced or appended to.
The 40-line orchestrator broke on the rejection edge — a graph makes that edge, and its bound, an explicit object.

What the graph layer actually provides

LangGraph models a multi-agent workflow as a directed graph. Nodes are functions — usually agents. Edges say what runs next. A single state object flows through, and each node returns a partial update to it.

Three properties follow, and each maps directly onto a failure of the hand-rolled version:

PropertyWhat it meansThe hand-rolled failure it removes
Explicit control flowRouting is a named function returning a node nameSix interacting boolean flags nobody can reason about
Durable stateState is checkpointed after every stepA kill at call 30 loses all 30 calls
Declared concurrencyFan-out is an edge topology, not thread codeSequential execution of independent agents

You could build all three yourself. Most teams that try end up with a worse version of exactly this, which is the honest argument for using the layer.

State is the message bus

This is the central idea and the one that determines whether your graph works. Agents in a LangGraph do not send each other messages. They read from and write to one shared state object. Agent B sees agent A's output because A wrote it into state.

Python
from typing import Annotatedfrom typing_extensions import TypedDictfrom operator import addclass ResearchState(TypedDict):    topic: str    sources: Annotated[list[dict], add]   # many nodes append here    draft: str                            # exactly one node writes here    critique: str    revision_count: int    approved: bool

The Annotated[list[dict], add] is not decoration. It is a reducer: a function saying how to combine a node's update with what is already in state. Without one, a write replaces the value. With add, the returned list is concatenated onto the existing list.

That distinction becomes critical the moment two nodes run in the same step. If two research agents both return {"sources": [...]} and sources has no reducer, LangGraph raises:

Text
InvalidUpdateError: At key 'sources': Can receive only one value per step.Use an Annotated key to handle multiple values.

That error is a gift. The alternative — which is what a hand-rolled dictionary gives you — is one agent's findings silently overwriting the other's, producing a report that is quietly missing half its evidence with no error anywhere.

A reducer is a written-down answer to "what happens when two agents write the same field at the same time". Fields without reducers are fields where that question has not been answered.

Choosing reducers

Field kindReducerWhy
Accumulating evidence, logs, messagesoperator.addEvery contribution must survive
A counteroperator.add on an intIncrements from concurrent nodes all count
Single-author artefact (the draft)None — last write winsOnly one node writes it; a reducer would hide a bug
Per-agent results keyed by agentA dict-merge reducerEach agent owns its own key, so no collision is possible
A boolean flag several nodes may setCustom, e.g. logical ORLast-write-wins on a flag is a race
Python
def merge_by_agent(left: dict, right: dict) -> dict:    """Each agent owns its own key. Collisions are a programming error."""    overlap = left.keys() & right.keys()    if overlap:        raise ValueError(f"two agents wrote the same key: {overlap}")    return {**left, **right}class State(TypedDict):    findings: Annotated[dict[str, dict], merge_by_agent]

That reducer raising on overlap is deliberate. Silent merging of two agents' results into one key is precisely the bug you cannot find later.

The minimal graph

Python
from langgraph.graph import StateGraph, START, ENDdef researcher(state: ResearchState) -> dict:    docs = search(state["topic"], k=8)    return {"sources": docs}              # appended, thanks to the reducerdef writer(state: ResearchState) -> dict:    text = model.generate(WRITE_PROMPT.format(        topic=state["topic"], sources=state["sources"],        critique=state.get("critique", "")))    return {"draft": text, "revision_count": state["revision_count"] + 1}def critic(state: ResearchState) -> dict:    verdict = model.generate(CRITIQUE_PROMPT.format(draft=state["draft"]))    return {"critique": verdict["notes"], "approved": verdict["approved"]}builder = StateGraph(ResearchState)builder.add_node("researcher", researcher)builder.add_node("writer", writer)builder.add_node("critic", critic)builder.add_edge(START, "researcher")builder.add_edge("researcher", "writer")builder.add_edge("writer", "critic")builder.add_edge("critic", END)graph = builder.compile()graph.invoke({"topic": "EU payment regulation", "sources": [],              "revision_count": 0, "approved": False})

Each node returns a partial update — only the keys it changed. It never mutates state in place. Mutating in place is the most common beginner error: it appears to work single-threaded and produces impossible-to-reproduce bugs the moment anything runs in parallel, because two nodes are then mutating the same object.

Parallel agent nodes

Add two edges out of one node and both targets run in the same superstep — LangGraph's unit of execution. Every node in a superstep sees the same input state, they run concurrently, and their updates are merged through the reducers before the next superstep begins.

Python
builder.add_node("academic", academic_agent)builder.add_node("news", news_agent)builder.add_node("filings", filings_agent)builder.add_node("synthesise", synthesiser)builder.add_edge(START, "academic")       # fan out: same source nodebuilder.add_edge(START, "news")builder.add_edge(START, "filings")builder.add_edge("academic", "synthesise")   # fan in: same target nodebuilder.add_edge("news", "synthesise")builder.add_edge("filings", "synthesise")

synthesise runs once, after all three finish. That is automatic: a node runs when all its incoming edges have delivered.

The speed-up, computed honestly

Say the three research agents take 4.1 s, 6.8 s and 3.2 s. Sequentially that is 4.1 + 6.8 + 3.2 = 14.1 s. In one superstep it is max(4.1, 6.8, 3.2) = 6.8 s — a 2.07× speed-up on that stage.

But the stage is not the whole pipeline. Add a 2.5 s planner before and a 1.8 s synthesiser after, both inherently sequential:

sequential=2.5+14.1+1.8=18.4 sparallel=2.5+6.8+1.8=11.1 s\text{sequential} = 2.5 + 14.1 + 1.8 = 18.4\ \text{s} \qquad \text{parallel} = 2.5 + 6.8 + 1.8 = 11.1\ \text{s}

That is 1.66×, not 2.07×. And the ceiling is fixed: even with infinitely fast research agents you cannot go below 4.3 s, so the maximum possible speed-up is 18.4 / 4.3 = 4.3×. Parallelising the research agents further — six instead of three — moves you toward 4.3× and never past it. Knowing that number stops you optimising the wrong stage.

Note also that the superstep is bounded by its slowest member. One agent taking 30 s makes the whole superstep 30 s no matter how fast the others are, so a per-node timeout matters more here than in sequential code.

Dynamic fan-out with Send

Static edges require you to know the number of parallel branches at build time. When the count depends on runtime data — one agent per source discovered — use Send.

Python
from langgraph.types import Senddef fan_out(state: State) -> list[Send]:    # One "analyse" node instance per source, created at runtime.    return [Send("analyse", {"source": s, "topic": state["topic"]})            for s in state["sources"]]builder.add_conditional_edges("gather", fan_out, ["analyse"])

Each Send carries its own private input to that node instance, so analyse receives one source rather than the whole list. Their returns merge through the state reducers exactly as static fan-out does. This is the map half of map-reduce, and the reducer is the reduce half.

Guard the count. len(state["sources"]) can be 400 after a good search, and 400 concurrent model calls will hit your rate limit and cost far more than they return. Slice explicitly: state["sources"][:12].

Conditional routing

This is where the six boolean flags go to die. A routing function reads state and returns the name of the next node.

Python
def route_after_critic(state: ResearchState) -> str:    if state["approved"]:        return "publish"    if state["revision_count"] >= 3:        return "escalate"                 # the guard that was missing before    return "writer"                       # loop back for revisionbuilder.add_conditional_edges(    "critic",    route_after_critic,    {"publish": "publish", "escalate": "escalate", "writer": "writer"},)

The loop that ran 47 times is now three lines, and its termination condition is a single visible comparison rather than an emergent property of flag assignments scattered across 220 lines. The third argument — the mapping from returned label to node — is optional when labels equal node names, but writing it out declares every possible destination up front, so the graph can be drawn and checked, and an unexpected label raises an error at run time instead of routing somewhere surprising.

Command: routing and updating together

Sometimes a node needs to both write state and decide where to go — the classic supervisor. Command does both in one return.

Python
from typing import Literalfrom langgraph.types import Commanddef supervisor(state: State) -> Command[Literal["research", "write", "escalate",                                               "__end__"]]:    decision = model.generate(SUPERVISOR_PROMPT.format(state=summarise(state)))    if decision["next"] == "done":        return Command(goto="__end__", update={"final": decision["answer"]})    return Command(goto=decision["next"],                   update={"instruction": decision["instruction"]})

The Literal in the return annotation is how LangGraph learns the possible destinations, so the graph can still be drawn and validated. Omit it and the compiled graph has no edges out of supervisor in its diagram, which makes visualisation useless exactly where it would help most.

The routing decision here comes from a model, which means it can return a node that does not exist. Validate:

Python
VALID = {"research", "write", "done"}if decision["next"] not in VALID:    return Command(goto="escalate",                   update={"error": f"invalid route {decision['next']!r}"})

Sub-graphs

A compiled graph is itself a valid node. That lets you build a research crew once and drop it into any workflow as a single box.

Python
# Inner graph: a three-agent research crewcrew_builder = StateGraph(CrewState)crew_builder.add_node("plan", plan_searches)crew_builder.add_node("search", run_searches)crew_builder.add_node("dedupe", dedupe_sources)crew_builder.add_edge(START, "plan")crew_builder.add_edge("plan", "search")crew_builder.add_edge("search", "dedupe")crew_builder.add_edge("dedupe", END)crew = crew_builder.compile()# Outer graph uses it as one nodeouter = StateGraph(ResearchState)outer.add_node("research_crew", crew)      # a graph, used as a nodeouter.add_node("writer", writer)outer.add_edge(START, "research_crew")outer.add_edge("research_crew", "writer")

The rule that matters: shared state keys are how parent and child communicate. If CrewState and ResearchState both declare sources, the crew's writes land in the parent's sources. If the crew declares docs instead, nothing crosses and the writer receives an empty list — with no error, because both graphs are individually valid.

When the schemas genuinely differ, wrap the sub-graph in a function that translates:

Python
def research_node(state: ResearchState) -> dict:    inner = crew.invoke({"query": state["topic"], "depth": 2})    # map in    return {"sources": inner["docs"]}                             # map outouter.add_node("research_crew", research_node)

That wrapper is three lines and it makes the interface explicit, which is worth more than the convenience of matching key names by accident.

Choosing the right primitive

You needUseNot
A always runs after Badd_edgeA conditional edge that always returns the same node
Choose among a fixed set of next nodesadd_conditional_edgesFlags in state read by the next node
Update state and choose the next nodeCommandA node followed by a router that re-reads what it wrote
A known, fixed number of parallel branchesMultiple add_edge callsSend, which adds needless indirection
A runtime-determined number of branchesSendA loop inside one node — that runs sequentially
Reuse a whole workflow in several placesA compiled sub-graphCopy-pasting nodes and edges
Several agents to write one fieldAn Annotated reducerHoping the ordering works out
Pause for a human decisioninterrupt() with a checkpointerA blocking input() inside a node

Durability, and why it is the real payoff

Compile with a checkpointer and the state is saved after every superstep, under a thread ID.

Python
from langgraph.checkpoint.memory import InMemorySavergraph = builder.compile(checkpointer=InMemorySaver())config = {"configurable": {"thread_id": "post-8821"}}graph.invoke({"topic": "EU payment regulation", "sources": [],              "revision_count": 0, "approved": False}, config)# Later, in a different process, after a crash:snapshot = graph.get_state(config)print(snapshot.values["revision_count"])   # 2print(snapshot.next)                       # ('critic',)graph.invoke(None, config)                 # resumes from the critic

Passing None as the input resumes rather than restarts. For a workflow that has already spent 30 model calls, that is the difference between losing the work and losing a second. Use InMemorySaver for development and a persistent saver — SQLite or Postgres — for anything real; in-memory checkpoints die with the process, which defeats the purpose.

Checkpointing turns an agent workflow from a function call into a resumable process. That single change is what makes long, expensive, multi-agent runs safe to operate.

Debugging graphs

Three techniques, in the order you should reach for them.

Stream the updates. Rather than waiting for a final answer, watch each node's contribution as it lands:

Python
for chunk in graph.stream(inputs, config, stream_mode="updates"):    for node, update in chunk.items():        print(f"{node:>14} -> {list(update.keys())}")# sub-graph internals too:for chunk in graph.stream(inputs, config, stream_mode="updates",                          subgraphs=True):    ...

This immediately answers the two most common questions: which node ran, and what did it actually write. A node appearing that you did not expect is a routing bug; a node writing keys you did not expect is a state-schema bug.

Read the state history. graph.get_state_history(config) returns every checkpoint, newest first, each with the state values and the next node. The revision loop that ran 47 times is unmistakable here: revision_count climbing while approved stays false and next alternates between two nodes.

Draw it. graph.get_graph().draw_mermaid() emits a diagram of the compiled topology. Diagrams do not catch logic errors, but they catch structural ones instantly — a node with no incoming edge, a conditional edge whose mapping omits a case, a sub-graph that is not connected to anything.

Named failure modes

Mutating state in place. A node does state["sources"].append(doc) and returns state. Reducers never fire, parallel nodes race on one list, and results vanish non-deterministically. Always return a new partial dict.

The missing reducer. Two parallel nodes write the same key. Best case, InvalidUpdateError. Worst case, the key is written by one node in one superstep and another in the next, so no error appears and the earlier value is quietly replaced.

The unbounded loop. A conditional edge that can return to an earlier node without a counter in state. This is the original 47-call bug reproduced in a graph. Every cycle needs a counter field and an explicit cap, and recursion_limit in the config is a backstop, not the design.

Fan-out into a rate limit. Send over an unbounded list. The graph is correct; your API returns 429 for forty of the eighty calls, and the retries cost more than the work. Cap the slice.

Sub-graph key mismatch. Parent and child use different names for the same concept. Nothing errors; downstream nodes just receive empty values. Prefer an explicit wrapper over relying on matching key names.

Assuming ordering within a superstep. Nodes in one superstep are concurrent; there is no guaranteed order between them. Code that works because "news always finishes before filings" will break the day a network call is slow.

What this means when you build

Design the state schema before any node. It is the interface between every agent in the system, and getting it wrong is expensive to change later because every node's reads and writes depend on it. For each field ask three questions in order: who writes this, can two writers collide, and what should happen if they do. The answer to the third is the reducer, and writing it down is the single highest-value thing you do in a LangGraph project.

Keep nodes pure with respect to state: read from the argument, return a partial dict, never mutate. That one discipline makes nodes unit-testable without the graph — you call the function with a dictionary and assert on the dictionary that comes back, no model, no runtime, no checkpointer.

Put every loop's counter in state, not in a closure or a module-level variable. State is what gets checkpointed, so a counter outside it resets on resume, and a retry loop whose counter resets on resume is an infinite loop with extra steps.

Finally, use a real checkpointer from the first day the workflow costs more than a few pennies to run. The instinct is to add durability later, once things are working. But "things are working" is measured on happy paths, and the run you most need to resume is the one that failed at step 12 of 15 after eleven minutes and a large model bill.