Course Content
Multi-Agent Systems and Collaboration
4 sections · 12 lessons
LangGraph for Multi-Agent Flows
A team hand-rolled a three-agent content pipeline: a researcher, a writer, and a critic. The orchestrator was a single Python function. It started at 40 lines. Then the product owner asked for one more behaviour — if the critic rejects the draft, send it back to the writer, up to three times.
That function grew to 220 lines and six boolean flags: has_research, has_draft, needs_revision, critic_ran, escalated, done. Then a bug appeared. needs_revision was set inside the retry branch but never cleared when the writer succeeded, so the loop ran until the step cap: 47 model calls for one blog post, roughly forty of them redundant.
The bug was trivial to fix and the shape of the code guaranteed it would come back. Control flow lived in flags, state lived in local variables, and there was no way to see what had happened except by reading logs. When the process was killed at call 30, everything was lost — there was nothing to resume from, because the state existed only in a stack frame.
That is the problem a graph orchestration layer solves, and it is worth being precise about what it actually gives you, because "you can draw it as a diagram" is not the reason.
What the graph layer actually provides
LangGraph models a multi-agent workflow as a directed graph. Nodes are functions — usually agents. Edges say what runs next. A single state object flows through, and each node returns a partial update to it.
Three properties follow, and each maps directly onto a failure of the hand-rolled version:
| Property | What it means | The hand-rolled failure it removes |
|---|---|---|
| Explicit control flow | Routing is a named function returning a node name | Six interacting boolean flags nobody can reason about |
| Durable state | State is checkpointed after every step | A kill at call 30 loses all 30 calls |
| Declared concurrency | Fan-out is an edge topology, not thread code | Sequential execution of independent agents |
You could build all three yourself. Most teams that try end up with a worse version of exactly this, which is the honest argument for using the layer.
State is the message bus
This is the central idea and the one that determines whether your graph works. Agents in a LangGraph do not send each other messages. They read from and write to one shared state object. Agent B sees agent A's output because A wrote it into state.
1from typing import Annotated2from typing_extensions import TypedDict3from operator import add45class ResearchState(TypedDict):6 topic: str7 sources: Annotated[list[dict], add] # many nodes append here8 draft: str # exactly one node writes here9 critique: str10 revision_count: int11 approved: boolThe Annotated[list[dict], add] is not decoration. It is a reducer: a function saying how to combine a node's update with what is already in state. Without one, a write replaces the value. With add, the returned list is concatenated onto the existing list.
That distinction becomes critical the moment two nodes run in the same step. If two research agents both return {"sources": [...]} and sources has no reducer, LangGraph raises:
InvalidUpdateError: At key 'sources': Can receive only one value per step.Use an Annotated key to handle multiple values.That error is a gift. The alternative — which is what a hand-rolled dictionary gives you — is one agent's findings silently overwriting the other's, producing a report that is quietly missing half its evidence with no error anywhere.
A reducer is a written-down answer to "what happens when two agents write the same field at the same time". Fields without reducers are fields where that question has not been answered.
Choosing reducers
| Field kind | Reducer | Why |
|---|---|---|
| Accumulating evidence, logs, messages | operator.add | Every contribution must survive |
| A counter | operator.add on an int | Increments from concurrent nodes all count |
| Single-author artefact (the draft) | None — last write wins | Only one node writes it; a reducer would hide a bug |
| Per-agent results keyed by agent | A dict-merge reducer | Each agent owns its own key, so no collision is possible |
| A boolean flag several nodes may set | Custom, e.g. logical OR | Last-write-wins on a flag is a race |
1def merge_by_agent(left: dict, right: dict) -> dict:2 """Each agent owns its own key. Collisions are a programming error."""3 overlap = left.keys() & right.keys()4 if overlap:5 raise ValueError(f"two agents wrote the same key: {overlap}")6 return {**left, **right}78class State(TypedDict):9 findings: Annotated[dict[str, dict], merge_by_agent]That reducer raising on overlap is deliberate. Silent merging of two agents' results into one key is precisely the bug you cannot find later.
The minimal graph
1from langgraph.graph import StateGraph, START, END23def researcher(state: ResearchState) -> dict:4 docs = search(state["topic"], k=8)5 return {"sources": docs} # appended, thanks to the reducer67def writer(state: ResearchState) -> dict:8 text = model.generate(WRITE_PROMPT.format(9 topic=state["topic"], sources=state["sources"],10 critique=state.get("critique", "")))11 return {"draft": text, "revision_count": state["revision_count"] + 1}1213def critic(state: ResearchState) -> dict:14 verdict = model.generate(CRITIQUE_PROMPT.format(draft=state["draft"]))15 return {"critique": verdict["notes"], "approved": verdict["approved"]}1617builder = StateGraph(ResearchState)18builder.add_node("researcher", researcher)19builder.add_node("writer", writer)20builder.add_node("critic", critic)21builder.add_edge(START, "researcher")22builder.add_edge("researcher", "writer")23builder.add_edge("writer", "critic")24builder.add_edge("critic", END)25graph = builder.compile()2627graph.invoke({"topic": "EU payment regulation", "sources": [],28 "revision_count": 0, "approved": False})Each node returns a partial update — only the keys it changed. It never mutates state in place. Mutating in place is the most common beginner error: it appears to work single-threaded and produces impossible-to-reproduce bugs the moment anything runs in parallel, because two nodes are then mutating the same object.
Parallel agent nodes
Add two edges out of one node and both targets run in the same superstep — LangGraph's unit of execution. Every node in a superstep sees the same input state, they run concurrently, and their updates are merged through the reducers before the next superstep begins.
1builder.add_node("academic", academic_agent)2builder.add_node("news", news_agent)3builder.add_node("filings", filings_agent)4builder.add_node("synthesise", synthesiser)56builder.add_edge(START, "academic") # fan out: same source node7builder.add_edge(START, "news")8builder.add_edge(START, "filings")910builder.add_edge("academic", "synthesise") # fan in: same target node11builder.add_edge("news", "synthesise")12builder.add_edge("filings", "synthesise")synthesise runs once, after all three finish. That is automatic: a node runs when all its incoming edges have delivered.
The speed-up, computed honestly
Say the three research agents take 4.1 s, 6.8 s and 3.2 s. Sequentially that is 4.1 + 6.8 + 3.2 = 14.1 s. In one superstep it is max(4.1, 6.8, 3.2) = 6.8 s — a 2.07× speed-up on that stage.
But the stage is not the whole pipeline. Add a 2.5 s planner before and a 1.8 s synthesiser after, both inherently sequential:
That is 1.66×, not 2.07×. And the ceiling is fixed: even with infinitely fast research agents you cannot go below 4.3 s, so the maximum possible speed-up is 18.4 / 4.3 = 4.3×. Parallelising the research agents further — six instead of three — moves you toward 4.3× and never past it. Knowing that number stops you optimising the wrong stage.
Note also that the superstep is bounded by its slowest member. One agent taking 30 s makes the whole superstep 30 s no matter how fast the others are, so a per-node timeout matters more here than in sequential code.
Dynamic fan-out with Send
Static edges require you to know the number of parallel branches at build time. When the count depends on runtime data — one agent per source discovered — use Send.
1from langgraph.types import Send23def fan_out(state: State) -> list[Send]:4 # One "analyse" node instance per source, created at runtime.5 return [Send("analyse", {"source": s, "topic": state["topic"]})6 for s in state["sources"]]78builder.add_conditional_edges("gather", fan_out, ["analyse"])Each Send carries its own private input to that node instance, so analyse receives one source rather than the whole list. Their returns merge through the state reducers exactly as static fan-out does. This is the map half of map-reduce, and the reducer is the reduce half.
Guard the count. len(state["sources"]) can be 400 after a good search, and 400 concurrent model calls will hit your rate limit and cost far more than they return. Slice explicitly: state["sources"][:12].
Conditional routing
This is where the six boolean flags go to die. A routing function reads state and returns the name of the next node.
1def route_after_critic(state: ResearchState) -> str:2 if state["approved"]:3 return "publish"4 if state["revision_count"] >= 3:5 return "escalate" # the guard that was missing before6 return "writer" # loop back for revision78builder.add_conditional_edges(9 "critic",10 route_after_critic,11 {"publish": "publish", "escalate": "escalate", "writer": "writer"},12)The loop that ran 47 times is now three lines, and its termination condition is a single visible comparison rather than an emergent property of flag assignments scattered across 220 lines. The third argument — the mapping from returned label to node — is optional when labels equal node names, but writing it out declares every possible destination up front, so the graph can be drawn and checked, and an unexpected label raises an error at run time instead of routing somewhere surprising.
Command: routing and updating together
Sometimes a node needs to both write state and decide where to go — the classic supervisor. Command does both in one return.
1from typing import Literal2from langgraph.types import Command34def supervisor(state: State) -> Command[Literal["research", "write", "escalate",5 "__end__"]]:6 decision = model.generate(SUPERVISOR_PROMPT.format(state=summarise(state)))7 if decision["next"] == "done":8 return Command(goto="__end__", update={"final": decision["answer"]})9 return Command(goto=decision["next"],10 update={"instruction": decision["instruction"]})The Literal in the return annotation is how LangGraph learns the possible destinations, so the graph can still be drawn and validated. Omit it and the compiled graph has no edges out of supervisor in its diagram, which makes visualisation useless exactly where it would help most.
The routing decision here comes from a model, which means it can return a node that does not exist. Validate:
1VALID = {"research", "write", "done"}2if decision["next"] not in VALID:3 return Command(goto="escalate",4 update={"error": f"invalid route {decision['next']!r}"})Sub-graphs
A compiled graph is itself a valid node. That lets you build a research crew once and drop it into any workflow as a single box.
1# Inner graph: a three-agent research crew2crew_builder = StateGraph(CrewState)3crew_builder.add_node("plan", plan_searches)4crew_builder.add_node("search", run_searches)5crew_builder.add_node("dedupe", dedupe_sources)6crew_builder.add_edge(START, "plan")7crew_builder.add_edge("plan", "search")8crew_builder.add_edge("search", "dedupe")9crew_builder.add_edge("dedupe", END)10crew = crew_builder.compile()1112# Outer graph uses it as one node13outer = StateGraph(ResearchState)14outer.add_node("research_crew", crew) # a graph, used as a node15outer.add_node("writer", writer)16outer.add_edge(START, "research_crew")17outer.add_edge("research_crew", "writer")The rule that matters: shared state keys are how parent and child communicate. If CrewState and ResearchState both declare sources, the crew's writes land in the parent's sources. If the crew declares docs instead, nothing crosses and the writer receives an empty list — with no error, because both graphs are individually valid.
When the schemas genuinely differ, wrap the sub-graph in a function that translates:
1def research_node(state: ResearchState) -> dict:2 inner = crew.invoke({"query": state["topic"], "depth": 2}) # map in3 return {"sources": inner["docs"]} # map out45outer.add_node("research_crew", research_node)That wrapper is three lines and it makes the interface explicit, which is worth more than the convenience of matching key names by accident.
Choosing the right primitive
| You need | Use | Not |
|---|---|---|
| A always runs after B | add_edge | A conditional edge that always returns the same node |
| Choose among a fixed set of next nodes | add_conditional_edges | Flags in state read by the next node |
| Update state and choose the next node | Command | A node followed by a router that re-reads what it wrote |
| A known, fixed number of parallel branches | Multiple add_edge calls | Send, which adds needless indirection |
| A runtime-determined number of branches | Send | A loop inside one node — that runs sequentially |
| Reuse a whole workflow in several places | A compiled sub-graph | Copy-pasting nodes and edges |
| Several agents to write one field | An Annotated reducer | Hoping the ordering works out |
| Pause for a human decision | interrupt() with a checkpointer | A blocking input() inside a node |
Durability, and why it is the real payoff
Compile with a checkpointer and the state is saved after every superstep, under a thread ID.
1from langgraph.checkpoint.memory import InMemorySaver23graph = builder.compile(checkpointer=InMemorySaver())4config = {"configurable": {"thread_id": "post-8821"}}56graph.invoke({"topic": "EU payment regulation", "sources": [],7 "revision_count": 0, "approved": False}, config)89# Later, in a different process, after a crash:10snapshot = graph.get_state(config)11print(snapshot.values["revision_count"]) # 212print(snapshot.next) # ('critic',)13graph.invoke(None, config) # resumes from the criticPassing None as the input resumes rather than restarts. For a workflow that has already spent 30 model calls, that is the difference between losing the work and losing a second. Use InMemorySaver for development and a persistent saver — SQLite or Postgres — for anything real; in-memory checkpoints die with the process, which defeats the purpose.
Checkpointing turns an agent workflow from a function call into a resumable process. That single change is what makes long, expensive, multi-agent runs safe to operate.
Debugging graphs
Three techniques, in the order you should reach for them.
Stream the updates. Rather than waiting for a final answer, watch each node's contribution as it lands:
1for chunk in graph.stream(inputs, config, stream_mode="updates"):2 for node, update in chunk.items():3 print(f"{node:>14} -> {list(update.keys())}")4# sub-graph internals too:5for chunk in graph.stream(inputs, config, stream_mode="updates",6 subgraphs=True):7 ...This immediately answers the two most common questions: which node ran, and what did it actually write. A node appearing that you did not expect is a routing bug; a node writing keys you did not expect is a state-schema bug.
Read the state history. graph.get_state_history(config) returns every checkpoint, newest first, each with the state values and the next node. The revision loop that ran 47 times is unmistakable here: revision_count climbing while approved stays false and next alternates between two nodes.
Draw it. graph.get_graph().draw_mermaid() emits a diagram of the compiled topology. Diagrams do not catch logic errors, but they catch structural ones instantly — a node with no incoming edge, a conditional edge whose mapping omits a case, a sub-graph that is not connected to anything.
Named failure modes
Mutating state in place. A node does state["sources"].append(doc) and returns state. Reducers never fire, parallel nodes race on one list, and results vanish non-deterministically. Always return a new partial dict.
The missing reducer. Two parallel nodes write the same key. Best case, InvalidUpdateError. Worst case, the key is written by one node in one superstep and another in the next, so no error appears and the earlier value is quietly replaced.
The unbounded loop. A conditional edge that can return to an earlier node without a counter in state. This is the original 47-call bug reproduced in a graph. Every cycle needs a counter field and an explicit cap, and recursion_limit in the config is a backstop, not the design.
Fan-out into a rate limit. Send over an unbounded list. The graph is correct; your API returns 429 for forty of the eighty calls, and the retries cost more than the work. Cap the slice.
Sub-graph key mismatch. Parent and child use different names for the same concept. Nothing errors; downstream nodes just receive empty values. Prefer an explicit wrapper over relying on matching key names.
Assuming ordering within a superstep. Nodes in one superstep are concurrent; there is no guaranteed order between them. Code that works because "news always finishes before filings" will break the day a network call is slow.
What this means when you build
Design the state schema before any node. It is the interface between every agent in the system, and getting it wrong is expensive to change later because every node's reads and writes depend on it. For each field ask three questions in order: who writes this, can two writers collide, and what should happen if they do. The answer to the third is the reducer, and writing it down is the single highest-value thing you do in a LangGraph project.
Keep nodes pure with respect to state: read from the argument, return a partial dict, never mutate. That one discipline makes nodes unit-testable without the graph — you call the function with a dictionary and assert on the dictionary that comes back, no model, no runtime, no checkpointer.
Put every loop's counter in state, not in a closure or a module-level variable. State is what gets checkpointed, so a counter outside it resets on resume, and a retry loop whose counter resets on resume is an infinite loop with extra steps.
Finally, use a real checkpointer from the first day the workflow costs more than a few pennies to run. The instinct is to add durability later, once things are working. But "things are working" is measured on happy paths, and the run you most need to resume is the one that failed at step 12 of 15 after eleven minutes and a large model bill.