Course Content
Multi-Agent Systems and Collaboration
4 sections · 12 lessons
Collaborative Research Assistants — The Capstone Build
A team set out to build a collaborative research assistant: give it a question, get back a cited report. They wrote five agents over a weekend — planner, searcher, verifier, synthesiser, critic — and wired them together on Sunday night.
The first end-to-end run cost 41 dollars and produced a 2,100-word report citing three sources that did not exist. Reading the logs took longer than writing the code. Three separate defects had combined:
- The critic could send the report back for revision with no cap. It did so 38 times, each cycle re-running the synthesiser. A properly guarded run costs well under a dollar; this one ran 38 revision cycles.
- The searcher wrote its results into
state["results"]; the verifier readstate["sources"]. Both graphs were individually valid. The verifier received an empty list every time, verified nothing, and reported success. - Because the verifier was a no-op, the synthesiser was free to invent citations, and nothing downstream checked them.
None of these was a hard bug. All three were consequences of building the agents before designing the system they live in. This is the build done in the other order: architecture, then state, then agents, then wiring, then the defences that keep it from costing 41 dollars.
Step 1: design the architecture before writing any code
Start from the requirements, stated concretely enough to argue about:
- Accept a research question in prose.
- Gather evidence from multiple independent sources.
- Verify every factual claim against a retrieved source before it appears in the report.
- Produce a report with inline citations.
- Finish within 3 minutes and 200,000 tokens.
- Return something useful even when parts fail.
Now derive the agents. The rule that matters: an agent boundary should follow a distinct set of tools and a distinct prompt, not a step in the workflow. Splitting "draft then polish" gives two agents with the same tools and nearly the same prompt, and you pay a handoff for nothing. Splitting "search the web" from "check a claim against a source" gives two agents with genuinely different tools and different success criteria.
| Agent | Reads | Writes | Tools | Model choice |
|---|---|---|---|---|
| Planner | question | subquestions | none | Small — this is decomposition, not reasoning over evidence |
| Searcher (× N) | one subquestion | sources | web search, fetch | Small — mostly tool orchestration |
| Verifier | sources | verified_claims | fetch | Strong — judgement about whether a source supports a claim |
| Synthesiser | verified_claims | draft | none | Strong — long-form writing |
| Critic | draft, verified_claims | critique, approved | none | Strong, different prompt from the synthesiser |
The critic reading verified_claims as well as the draft is deliberate. A critic that sees only the draft can judge style. A critic that also sees the evidence can catch the failure that cost the team their credibility: a sentence in the report with no supporting claim behind it.
┌──────────┐ question ─────►│ planner │ └─────┬────┘ ┌──────────┼──────────┐ (one searcher per subquestion, ┌────▼───┐ ┌────▼───┐ ┌────▼───┐ created at runtime) │search 1│ │search 2│ │search 3│ └────┬───┘ └────┬───┘ └────┬───┘ └──────────┼──────────┘ ┌─────▼────┐ │ verifier │ └─────┬────┘ ┌─────▼──────┐ │synthesiser │◄──────┐ └─────┬──────┘ │ revise (max 2) ┌─────▼────┐ │ │ critic ├─────────┘ └─────┬────┘ ▼ approved or revisions exhausted reportDraw the graph and label every edge with the state key it carries before you write a single agent. Most multi-agent bugs are edges nobody drew.
Step 2: define the shared state — and keep it consistent
The state schema is the contract between all five agents. The results versus sources mismatch happened because there was no schema; each agent invented a key when it needed one.
1from typing import Annotated, TypedDict2from dataclasses import dataclass, field, asdict3import hashlib, operator45@dataclass(frozen=True)6class Source:7 url: str8 title: str9 text: str10 fetched_at: float11 subquestion: str1213 @property14 def id(self) -> str:15 return hashlib.sha256(self.url.encode()).hexdigest()[:16]1617@dataclass(frozen=True)18class Claim:19 text: str20 source_id: str # MUST refer to a Source that exists21 quote: str # the exact supporting span from that source22 confidence: float # 0.0 - 1.02324def merge_sources(left: list[dict], right: list[dict]) -> list[dict]:25 """Concatenate and de-duplicate by URL hash, keeping first occurrence."""26 seen, out = set(), []27 for s in [*left, *right]:28 key = hashlib.sha256(s["url"].encode()).hexdigest()[:16]29 if key not in seen:30 seen.add(key)31 out.append(s)32 return out3334class ResearchState(TypedDict):35 question: str36 subquestions: list[str]37 sources: Annotated[list[dict], merge_sources] # many writers38 verified_claims: Annotated[list[dict], operator.add]39 draft: str # one writer40 critique: str41 approved: bool42 revisions: int43 tokens_used: Annotated[int, operator.add]44 errors: Annotated[list[dict], operator.add] # never raise, recordEvery reducer choice here is a decision about concurrency. sources has three or more concurrent writers, so it needs a merge — and the merge de-duplicates, because two searchers asking related subquestions routinely find the same page. Without de-duplication a popular page appears four times, the verifier checks it four times, and the token budget evaporates on repeated work.
draft deliberately has no reducer. Exactly one node writes it, and if a second one ever does, the resulting InvalidUpdateError is information rather than a silent overwrite.
tokens_used with operator.add means every node's spend accumulates automatically, including from parallel searchers. That single field is what makes the budget guard possible.
errors as an accumulating list encodes the sixth requirement. Nodes record failures into state instead of raising, so a searcher that dies does not take the report down with it.
Step 3: build each specialised agent
The planner
1PLAN_PROMPT = """Break this research question into 3-5 independent2sub-questions that can each be answered by a separate web search.3Return ONLY a JSON array of strings. No commentary.45QUESTION: {question}"""67def planner(state: ResearchState) -> dict:8 raw = small_model.generate(PLAN_PROMPT.format(question=state["question"]))9 try:10 subs = json.loads(extract_json(raw))11 assert isinstance(subs, list) and all(isinstance(s, str) for s in subs)12 subs = [s.strip() for s in subs if s.strip()][:5] # hard cap13 except (json.JSONDecodeError, AssertionError) as exc:14 # Degrade, do not fail: one search on the original question.15 return {"subquestions": [state["question"]],16 "errors": [{"node": "planner", "error": str(exc)}],17 "tokens_used": raw.usage.total}18 if not subs:19 subs = [state["question"]]20 return {"subquestions": subs, "tokens_used": raw.usage.total}The [:5] is not paranoia. Ask a model for "3-5" and it will occasionally return eleven, and since each sub-question spawns a searcher, eleven sub-questions is eleven concurrent agents and roughly double the budget. Caps belong in code, never in prompts alone.
The searchers, fanned out at runtime
1from langgraph.types import Send23def fan_out_searchers(state: ResearchState) -> list[Send]:4 return [Send("search", {"subquestion": q, "question": state["question"]})5 for q in state["subquestions"]]67def search(payload: dict) -> dict:8 q = payload["subquestion"]9 try:10 hits = web_search(q, k=4)11 sources = []12 for h in hits[:4]:13 text = fetch(h["url"], timeout=8, max_chars=20_000)14 if not text:15 continue16 sources.append(asdict(Source(url=h["url"], title=h["title"],17 text=text, fetched_at=time.time(),18 subquestion=q)))19 return {"sources": sources, "tokens_used": 0}20 except Exception as exc:21 return {"sources": [],22 "errors": [{"node": "search", "subquestion": q,23 "error": f"{type(exc).__name__}: {exc}"}]}Note max_chars=20_000. A single long page can be 400,000 characters, roughly 100,000 tokens — half the entire budget for one source. Truncate at fetch time, not at prompt time, so the oversized text never enters state at all.
The verifier
This is the agent that makes the report trustworthy, and the one the failed build skipped.
1VERIFY_PROMPT = """From the SOURCE below, extract factual claims that2directly answer the QUESTION. For each claim give the exact supporting3quote from the source. Do not infer. If the source does not answer the4question, return an empty array.56Return ONLY JSON: [{{"text": "...", "quote": "...", "confidence": 0.0-1.0}}]78QUESTION: {question}9SOURCE ({url}):10{text}"""1112def verifier(state: ResearchState) -> dict:13 claims, errors, tokens = [], [], 014 for s in state["sources"][:12]: # cap the work15 src = Source(**s)16 try:17 raw = strong_model.generate(VERIFY_PROMPT.format(18 question=state["question"], url=src.url,19 text=src.text[:12_000]))20 tokens += raw.usage.total21 for c in json.loads(extract_json(raw)):22 # The quote must genuinely appear in the source text.23 if c["quote"] and c["quote"][:120] in src.text:24 claims.append(asdict(Claim(text=c["text"],25 source_id=src.id,26 quote=c["quote"],27 confidence=float(c["confidence"]))))28 else:29 errors.append({"node": "verifier", "url": src.url,30 "error": "quote not found in source"})31 except Exception as exc:32 errors.append({"node": "verifier", "url": src.url,33 "error": f"{type(exc).__name__}: {exc}"})34 return {"verified_claims": claims, "errors": errors, "tokens_used": tokens}The substring check c["quote"][:120] in src.text is the single most valuable line in this build. A model asked to quote a source will sometimes produce a fluent, plausible quote that is not in the text. Checking mechanically that the quote actually appears converts an unverifiable assertion into a verifiable one, for the cost of a string comparison. In practice this rejects somewhere between 3% and 10% of extracted claims, and those are exactly the claims that would have become fake citations.
Do not ask a model to confirm that a model was truthful. Where a mechanical check is possible — does this quote appear in this text — use the mechanical check.
The synthesiser
1SYNTH_PROMPT = """Write a report answering the QUESTION using ONLY the2CLAIMS below. Every factual sentence must end with a citation marker3[source_id]. Do not state anything not present in the claims.4{revision_note}56QUESTION: {question}7CLAIMS:8{claims}"""910def synthesiser(state: ResearchState) -> dict:11 if not state["verified_claims"]:12 return {"draft": "", "approved": False,13 "errors": [{"node": "synthesiser",14 "error": "no verified claims to write from"}]}15 note = (f"Revise per this critique: {state['critique']}"16 if state.get("critique") else "")17 claims_block = "\n".join(18 f"[{c['source_id']}] {c['text']} (conf {c['confidence']:.2f})"19 for c in state["verified_claims"])20 raw = strong_model.generate(SYNTH_PROMPT.format(21 question=state["question"], claims=claims_block, revision_note=note))22 is_revision = bool(state.get("critique")) # the first draft is not one23 return {"draft": raw.text,24 "revisions": state["revisions"] + (1 if is_revision else 0),25 "tokens_used": raw.usage.total}The critic
1import re23def critic(state: ResearchState) -> dict:4 draft = state["draft"]5 valid_ids = {c["source_id"] for c in state["verified_claims"]}67 # Mechanical checks first — they are free and unambiguous.8 cited = set(re.findall(r"\[([0-9a-f]{16})\]", draft))9 fabricated = cited - valid_ids10 sentences = [s for s in re.split(r"(?<=[.!?])\s+", draft) if len(s) > 40]11 uncited = [s for s in sentences if not re.search(r"\[[0-9a-f]{16}\]", s)]1213 if fabricated:14 return {"approved": False,15 "critique": f"Citations not in evidence: {sorted(fabricated)}"}16 if len(uncited) > len(sentences) * 0.25:17 return {"approved": False,18 "critique": f"{len(uncited)} of {len(sentences)} sentences "19 f"lack a citation."}2021 # Only then spend tokens on a judgement call.22 raw = strong_model.generate(CRITIC_PROMPT.format(23 question=state["question"], draft=draft))24 verdict = json.loads(extract_json(raw))25 return {"approved": bool(verdict["approved"]),26 "critique": verdict.get("notes", ""),27 "tokens_used": raw.usage.total}Cheap checks before expensive ones. A fabricated citation is detectable with a set difference costing nothing, so there is no reason to spend 12,000 tokens asking a model about a draft that has already failed.
Step 4: wire the agents into the workflow
1import sqlite32from langgraph.graph import StateGraph, START, END3from langgraph.checkpoint.sqlite import SqliteSaver # langgraph-checkpoint-sqlite45MAX_REVISIONS = 26TOKEN_BUDGET = 200_00078def route_after_critic(state: ResearchState) -> str:9 if state["approved"]:10 return "finish"11 if state["revisions"] >= MAX_REVISIONS:12 return "finish" # ship the best draft we have13 if state["tokens_used"] > TOKEN_BUDGET * 0.85:14 return "finish" # no budget for another cycle15 return "synthesiser"1617def route_after_verifier(state: ResearchState) -> str:18 if not state["verified_claims"]:19 return "finish" # nothing to write; do not try20 return "synthesiser"2122builder = StateGraph(ResearchState)23builder.add_node("planner", planner)24builder.add_node("search", search)25builder.add_node("verifier", verifier)26builder.add_node("synthesiser", synthesiser)27builder.add_node("critic", critic)28builder.add_node("finish", finish)2930builder.add_edge(START, "planner")31builder.add_conditional_edges("planner", fan_out_searchers, ["search"])32builder.add_edge("search", "verifier")33builder.add_conditional_edges("verifier", route_after_verifier,34 {"synthesiser": "synthesiser", "finish": "finish"})35builder.add_edge("synthesiser", "critic")36builder.add_conditional_edges("critic", route_after_critic,37 {"synthesiser": "synthesiser", "finish": "finish"})38builder.add_edge("finish", END)3940saver = SqliteSaver(sqlite3.connect("research.db", check_same_thread=False))41graph = builder.compile(checkpointer=saver)Three guards live in route_after_critic and each closes one of the failure paths from the disastrous first run: a revision cap, a budget cap, and an approval condition. The 38-cycle loop cannot happen, because the second condition fires at two.
The finish node is not decoration either. It is where the report is assembled with its bibliography, where partial results are labelled as partial, and where the errors list becomes something a reader can see:
1def finish(state: ResearchState) -> dict:2 if not state["draft"]:3 return {"draft": ""} # nothing verified, so no report — never invent one4 by_id = {}5 for s in state["sources"]:6 by_id[Source(**s).id] = s7 used = sorted({c["source_id"] for c in state["verified_claims"]})8 biblio = "\n".join(f"[{i}] {by_id[i]['title']} — {by_id[i]['url']}"9 for i in used if i in by_id)10 caveats = []11 if state["errors"]:12 caveats.append(f"{len(state['errors'])} step(s) failed; "13 "coverage may be incomplete.")14 if not state["approved"]:15 caveats.append("Not approved by review; revisions exhausted.")16 return {"draft": state["draft"] + "\n\n## Sources\n" + biblio +17 ("\n\n> " + " ".join(caveats) if caveats else "")}Running it
1config = {"configurable": {"thread_id": "eu-payments-2026"}}2initial = {"question": "How has EU payment regulation changed since 2023?",3 "subquestions": [], "sources": [], "verified_claims": [],4 "draft": "", "critique": "", "approved": False,5 "revisions": 0, "tokens_used": 0, "errors": []}67for chunk in graph.stream(initial, config, stream_mode="updates"):8 for node, update in chunk.items():9 print(f"{node:>13} -> {list(update)}") planner -> ['subquestions', 'tokens_used'] search -> ['sources', 'tokens_used'] (×4, concurrent) verifier -> ['verified_claims', 'errors', 'tokens_used'] synthesiser -> ['draft', 'revisions', 'tokens_used'] critic -> ['approved', 'critique'] approved=False synthesiser -> ['draft', 'revisions', 'tokens_used'] critic -> ['approved', 'critique'] approved=True finish -> ['draft']A representative run: planner 2,100 tokens; four searchers, no model tokens but 11 pages fetched; verifier 11 sources × ~2,800 tokens = 30,800; synthesiser 24,600; critic 11,900; one revision cycle adds 24,600 + 11,900. Total 2,100 + 30,800 + 24,600 + 11,900 + 24,600 + 11,900 = 105,900 tokens, comfortably inside 200,000.
At example rates of roughly 3 dollars per million input tokens and 15 per million output (check your model's current prices), with about a 70/30 split, that is (0.7 × 105,900 × 3 + 0.3 × 105,900 × 15) / 1,000,000 = 0.22 + 0.48 = 0.70 dollars per report. The first, unguarded run cost 41 dollars — about 59 times as much — and the only difference is three comparisons in a routing function.
Wall-clock: planner 1.9 s, four searchers concurrently at max(3.1, 4.4, 2.8, 6.2) = 6.2 s, verifier 14.0 s, synthesiser 8.1 s, critic 3.3 s, revision cycle 11.4 s, finish 0.1 s — about 45 seconds, inside the 3-minute requirement. Had the searchers run sequentially they would have taken 3.1 + 4.4 + 2.8 + 6.2 = 16.5 s instead of 6.2, adding 10.3 seconds for nothing.
Step 5: defensive error handling
The governing principle: a node should return an error into state rather than raise, unless the error means no useful output is possible. A raised exception in one searcher aborts the whole graph and throws away the three searches that succeeded.
1import functools, time23def resilient(node_name: str, fallback: dict):4 """Wrap a node so failures become state, not exceptions."""5 def decorator(fn):6 @functools.wraps(fn)7 def wrapper(state):8 t0 = time.time()9 try:10 return fn(state)11 except Exception as exc:12 return {**fallback,13 "errors": [{"node": node_name,14 "error": f"{type(exc).__name__}: {exc}",15 "elapsed_s": round(time.time() - t0, 2)}]}16 return wrapper17 return decorator1819verifier = resilient("verifier", {"verified_claims": []})(verifier)20critic = resilient("critic", {"approved": True, "critique": ""})(critic)Choosing each fallback is a judgement about which way to fail. The verifier's fallback is an empty claim list, which routes to finish and produces no report — correct, because a report with unverified claims is worse than no report. The critic's fallback is approved: True, which ships the draft unreviewed — correct, because the mechanical citation checks have already run and a broken critic should not block delivery indefinitely.
| Failure | Response | User-visible outcome |
|---|---|---|
| One searcher throws | Record error, continue with other results | Report with a coverage caveat |
| All searchers return nothing | Route to finish | "No sources found", not a fabricated report |
| Model returns non-JSON | Catch, degrade to a safe default | Fewer sub-questions, still a report |
| Quote not found in source | Drop the claim, record the error | Report without that unverifiable claim |
| Critic never approves | Revision cap routes to finish | Best available draft, marked unapproved |
| Budget 85% consumed | Stop revising, finish | Report inside cost limits |
| Process killed mid-run | Checkpoint resume | graph.invoke(None, config) continues |
Two more defences worth adding before this touches real users. A per-node timeout, because a hung fetch inside a searcher blocks the superstep and therefore the whole graph — wrap the tool call, not the node, so the searcher can return partial results. And a circuit breaker on the search tool: after five consecutive failures, stop calling it for sixty seconds and let searchers return empty immediately, so that an outage at the search provider produces fast degraded reports rather than a slow queue of timeouts.
Where people get it wrong
Building agents before the state schema. This produced the results/sources mismatch and it produces one of these in nearly every first build. Write the TypedDict first and make every agent's signature reference it.
Trusting model JSON. Every json.loads on model output needs a try/except and a schema assertion. The failure is not rare and it is not graceful.
No cap on fan-out. Sub-question count, sources per search, sources sent to the verifier: all three need explicit numeric limits. A prompt saying "3 to 5" is a suggestion.
Letting the critic be the only stopping condition. An LLM critic that can always find something to improve will always find something to improve. Pair it with a hard revision cap and a budget cap.
Reporting partial success as success. Three of four searchers failing should be visible in the output, not just in the logs. The caveat line in finish costs three lines and prevents a reader treating a thin report as a thorough one.
What this means when you build
The five agents in this system are, individually, unremarkable — each one is a prompt, a tool or two, and some parsing. Nearly all the engineering value sits in four places that are not the agents at all: the state schema with its reducers, the routing functions with their caps, the mechanical verification of quotes, and the assembly step that tells the truth about what failed.
That ratio holds generally. When a multi-agent build goes wrong, the cause is almost never that an agent reasoned badly. It is that two agents disagreed about a key name, or a loop had no counter, or nobody checked something checkable, or a partial result was presented as a complete one.
So build in this order and resist the urge to reorder it: draw the graph and label the edges, write the state schema and choose every reducer deliberately, write the guards, and only then write the agents. The agents are the part you can iterate on cheaply — a prompt change is a minute's work. The state schema is the part that is expensive to change once five nodes depend on it, and the guards are the part whose absence you discover through a 41-dollar invoice.