Multi-Agent Systems and Collaboration

Collaborative Research Assistants — The Capstone Build


A team set out to build a collaborative research assistant: give it a question, get back a cited report. They wrote five agents over a weekend — planner, searcher, verifier, synthesiser, critic — and wired them together on Sunday night.

The first end-to-end run cost 41 dollars and produced a 2,100-word report citing three sources that did not exist. Reading the logs took longer than writing the code. Three separate defects had combined:

  • The critic could send the report back for revision with no cap. It did so 38 times, each cycle re-running the synthesiser. A properly guarded run costs well under a dollar; this one ran 38 revision cycles.
  • The searcher wrote its results into state["results"]; the verifier read state["sources"]. Both graphs were individually valid. The verifier received an empty list every time, verified nothing, and reported success.
  • Because the verifier was a no-op, the synthesiser was free to invent citations, and nothing downstream checked them.

None of these was a hard bug. All three were consequences of building the agents before designing the system they live in. This is the build done in the other order: architecture, then state, then agents, then wiring, then the defences that keep it from costing 41 dollars.

Five agents over one shared statePlanner splitsthe questionSearchers fannedout at run timeVerifierchecks eachclaim's sourceSynthesiserwrites thereportCritic accepts,or sends it backFan-out width is decided from the plan, so the graph cannot be wired with a fixed number of searchers.
The architecture is the shared state schema — five agents that disagree about a field name fail as one system.

Step 1: design the architecture before writing any code

Start from the requirements, stated concretely enough to argue about:

  1. Accept a research question in prose.
  2. Gather evidence from multiple independent sources.
  3. Verify every factual claim against a retrieved source before it appears in the report.
  4. Produce a report with inline citations.
  5. Finish within 3 minutes and 200,000 tokens.
  6. Return something useful even when parts fail.

Now derive the agents. The rule that matters: an agent boundary should follow a distinct set of tools and a distinct prompt, not a step in the workflow. Splitting "draft then polish" gives two agents with the same tools and nearly the same prompt, and you pay a handoff for nothing. Splitting "search the web" from "check a claim against a source" gives two agents with genuinely different tools and different success criteria.

AgentReadsWritesToolsModel choice
PlannerquestionsubquestionsnoneSmall — this is decomposition, not reasoning over evidence
Searcher (× N)one subquestionsourcesweb search, fetchSmall — mostly tool orchestration
Verifiersourcesverified_claimsfetchStrong — judgement about whether a source supports a claim
Synthesiserverified_claimsdraftnoneStrong — long-form writing
Criticdraft, verified_claimscritique, approvednoneStrong, different prompt from the synthesiser

The critic reading verified_claims as well as the draft is deliberate. A critic that sees only the draft can judge style. A critic that also sees the evidence can catch the failure that cost the team their credibility: a sentence in the report with no supporting claim behind it.

Text
                  ┌──────────┐   question ─────►│ planner  │                  └─────┬────┘             ┌──────────┼──────────┐        (one searcher per subquestion,        ┌────▼───┐ ┌────▼───┐ ┌────▼───┐     created at runtime)        │search 1│ │search 2│ │search 3│        └────┬───┘ └────┬───┘ └────┬───┘             └──────────┼──────────┘                  ┌─────▼────┐                  │ verifier │                  └─────┬────┘                  ┌─────▼──────┐                  │synthesiser │◄──────┐                  └─────┬──────┘       │ revise (max 2)                  ┌─────▼────┐         │                  │  critic  ├─────────┘                  └─────┬────┘                        ▼  approved or revisions exhausted                      report

Draw the graph and label every edge with the state key it carries before you write a single agent. Most multi-agent bugs are edges nobody drew.

Step 2: define the shared state — and keep it consistent

The state schema is the contract between all five agents. The results versus sources mismatch happened because there was no schema; each agent invented a key when it needed one.

Python
from typing import Annotated, TypedDictfrom dataclasses import dataclass, field, asdictimport hashlib, operator@dataclass(frozen=True)class Source:    url: str    title: str    text: str    fetched_at: float    subquestion: str    @property    def id(self) -> str:        return hashlib.sha256(self.url.encode()).hexdigest()[:16]@dataclass(frozen=True)class Claim:    text: str    source_id: str          # MUST refer to a Source that exists    quote: str              # the exact supporting span from that source    confidence: float       # 0.0 - 1.0def merge_sources(left: list[dict], right: list[dict]) -> list[dict]:    """Concatenate and de-duplicate by URL hash, keeping first occurrence."""    seen, out = set(), []    for s in [*left, *right]:        key = hashlib.sha256(s["url"].encode()).hexdigest()[:16]        if key not in seen:            seen.add(key)            out.append(s)    return outclass ResearchState(TypedDict):    question: str    subquestions: list[str]    sources: Annotated[list[dict], merge_sources]      # many writers    verified_claims: Annotated[list[dict], operator.add]    draft: str                                          # one writer    critique: str    approved: bool    revisions: int    tokens_used: Annotated[int, operator.add]    errors: Annotated[list[dict], operator.add]         # never raise, record

Every reducer choice here is a decision about concurrency. sources has three or more concurrent writers, so it needs a merge — and the merge de-duplicates, because two searchers asking related subquestions routinely find the same page. Without de-duplication a popular page appears four times, the verifier checks it four times, and the token budget evaporates on repeated work.

draft deliberately has no reducer. Exactly one node writes it, and if a second one ever does, the resulting InvalidUpdateError is information rather than a silent overwrite.

tokens_used with operator.add means every node's spend accumulates automatically, including from parallel searchers. That single field is what makes the budget guard possible.

errors as an accumulating list encodes the sixth requirement. Nodes record failures into state instead of raising, so a searcher that dies does not take the report down with it.

Step 3: build each specialised agent

The planner

Python
PLAN_PROMPT = """Break this research question into 3-5 independentsub-questions that can each be answered by a separate web search.Return ONLY a JSON array of strings. No commentary.QUESTION: {question}"""def planner(state: ResearchState) -> dict:    raw = small_model.generate(PLAN_PROMPT.format(question=state["question"]))    try:        subs = json.loads(extract_json(raw))        assert isinstance(subs, list) and all(isinstance(s, str) for s in subs)        subs = [s.strip() for s in subs if s.strip()][:5]     # hard cap    except (json.JSONDecodeError, AssertionError) as exc:        # Degrade, do not fail: one search on the original question.        return {"subquestions": [state["question"]],                "errors": [{"node": "planner", "error": str(exc)}],                "tokens_used": raw.usage.total}    if not subs:        subs = [state["question"]]    return {"subquestions": subs, "tokens_used": raw.usage.total}

The [:5] is not paranoia. Ask a model for "3-5" and it will occasionally return eleven, and since each sub-question spawns a searcher, eleven sub-questions is eleven concurrent agents and roughly double the budget. Caps belong in code, never in prompts alone.

The searchers, fanned out at runtime

Python
from langgraph.types import Senddef fan_out_searchers(state: ResearchState) -> list[Send]:    return [Send("search", {"subquestion": q, "question": state["question"]})            for q in state["subquestions"]]def search(payload: dict) -> dict:    q = payload["subquestion"]    try:        hits = web_search(q, k=4)        sources = []        for h in hits[:4]:            text = fetch(h["url"], timeout=8, max_chars=20_000)            if not text:                continue            sources.append(asdict(Source(url=h["url"], title=h["title"],                                         text=text, fetched_at=time.time(),                                         subquestion=q)))        return {"sources": sources, "tokens_used": 0}    except Exception as exc:        return {"sources": [],                "errors": [{"node": "search", "subquestion": q,                            "error": f"{type(exc).__name__}: {exc}"}]}

Note max_chars=20_000. A single long page can be 400,000 characters, roughly 100,000 tokens — half the entire budget for one source. Truncate at fetch time, not at prompt time, so the oversized text never enters state at all.

The verifier

This is the agent that makes the report trustworthy, and the one the failed build skipped.

Python
VERIFY_PROMPT = """From the SOURCE below, extract factual claims thatdirectly answer the QUESTION. For each claim give the exact supportingquote from the source. Do not infer. If the source does not answer thequestion, return an empty array.Return ONLY JSON: [{{"text": "...", "quote": "...", "confidence": 0.0-1.0}}]QUESTION: {question}SOURCE ({url}):{text}"""def verifier(state: ResearchState) -> dict:    claims, errors, tokens = [], [], 0    for s in state["sources"][:12]:                 # cap the work        src = Source(**s)        try:            raw = strong_model.generate(VERIFY_PROMPT.format(                question=state["question"], url=src.url,                text=src.text[:12_000]))            tokens += raw.usage.total            for c in json.loads(extract_json(raw)):                # The quote must genuinely appear in the source text.                if c["quote"] and c["quote"][:120] in src.text:                    claims.append(asdict(Claim(text=c["text"],                                               source_id=src.id,                                               quote=c["quote"],                                               confidence=float(c["confidence"]))))                else:                    errors.append({"node": "verifier", "url": src.url,                                   "error": "quote not found in source"})        except Exception as exc:            errors.append({"node": "verifier", "url": src.url,                           "error": f"{type(exc).__name__}: {exc}"})    return {"verified_claims": claims, "errors": errors, "tokens_used": tokens}

The substring check c["quote"][:120] in src.text is the single most valuable line in this build. A model asked to quote a source will sometimes produce a fluent, plausible quote that is not in the text. Checking mechanically that the quote actually appears converts an unverifiable assertion into a verifiable one, for the cost of a string comparison. In practice this rejects somewhere between 3% and 10% of extracted claims, and those are exactly the claims that would have become fake citations.

Do not ask a model to confirm that a model was truthful. Where a mechanical check is possible — does this quote appear in this text — use the mechanical check.

The synthesiser

Python
SYNTH_PROMPT = """Write a report answering the QUESTION using ONLY theCLAIMS below. Every factual sentence must end with a citation marker[source_id]. Do not state anything not present in the claims.{revision_note}QUESTION: {question}CLAIMS:{claims}"""def synthesiser(state: ResearchState) -> dict:    if not state["verified_claims"]:        return {"draft": "", "approved": False,                "errors": [{"node": "synthesiser",                            "error": "no verified claims to write from"}]}    note = (f"Revise per this critique: {state['critique']}"            if state.get("critique") else "")    claims_block = "\n".join(        f"[{c['source_id']}] {c['text']}  (conf {c['confidence']:.2f})"        for c in state["verified_claims"])    raw = strong_model.generate(SYNTH_PROMPT.format(        question=state["question"], claims=claims_block, revision_note=note))    is_revision = bool(state.get("critique"))     # the first draft is not one    return {"draft": raw.text,            "revisions": state["revisions"] + (1 if is_revision else 0),            "tokens_used": raw.usage.total}

The critic

Python
import redef critic(state: ResearchState) -> dict:    draft = state["draft"]    valid_ids = {c["source_id"] for c in state["verified_claims"]}    # Mechanical checks first — they are free and unambiguous.    cited = set(re.findall(r"\[([0-9a-f]{16})\]", draft))    fabricated = cited - valid_ids    sentences = [s for s in re.split(r"(?<=[.!?])\s+", draft) if len(s) > 40]    uncited = [s for s in sentences if not re.search(r"\[[0-9a-f]{16}\]", s)]    if fabricated:        return {"approved": False,                "critique": f"Citations not in evidence: {sorted(fabricated)}"}    if len(uncited) > len(sentences) * 0.25:        return {"approved": False,                "critique": f"{len(uncited)} of {len(sentences)} sentences "                            f"lack a citation."}    # Only then spend tokens on a judgement call.    raw = strong_model.generate(CRITIC_PROMPT.format(        question=state["question"], draft=draft))    verdict = json.loads(extract_json(raw))    return {"approved": bool(verdict["approved"]),            "critique": verdict.get("notes", ""),            "tokens_used": raw.usage.total}

Cheap checks before expensive ones. A fabricated citation is detectable with a set difference costing nothing, so there is no reason to spend 12,000 tokens asking a model about a draft that has already failed.

Step 4: wire the agents into the workflow

Python
import sqlite3from langgraph.graph import StateGraph, START, ENDfrom langgraph.checkpoint.sqlite import SqliteSaver   # langgraph-checkpoint-sqliteMAX_REVISIONS = 2TOKEN_BUDGET  = 200_000def route_after_critic(state: ResearchState) -> str:    if state["approved"]:        return "finish"    if state["revisions"] >= MAX_REVISIONS:        return "finish"                       # ship the best draft we have    if state["tokens_used"] > TOKEN_BUDGET * 0.85:        return "finish"                       # no budget for another cycle    return "synthesiser"def route_after_verifier(state: ResearchState) -> str:    if not state["verified_claims"]:        return "finish"                       # nothing to write; do not try    return "synthesiser"builder = StateGraph(ResearchState)builder.add_node("planner", planner)builder.add_node("search", search)builder.add_node("verifier", verifier)builder.add_node("synthesiser", synthesiser)builder.add_node("critic", critic)builder.add_node("finish", finish)builder.add_edge(START, "planner")builder.add_conditional_edges("planner", fan_out_searchers, ["search"])builder.add_edge("search", "verifier")builder.add_conditional_edges("verifier", route_after_verifier,                              {"synthesiser": "synthesiser", "finish": "finish"})builder.add_edge("synthesiser", "critic")builder.add_conditional_edges("critic", route_after_critic,                              {"synthesiser": "synthesiser", "finish": "finish"})builder.add_edge("finish", END)saver = SqliteSaver(sqlite3.connect("research.db", check_same_thread=False))graph = builder.compile(checkpointer=saver)

Three guards live in route_after_critic and each closes one of the failure paths from the disastrous first run: a revision cap, a budget cap, and an approval condition. The 38-cycle loop cannot happen, because the second condition fires at two.

The finish node is not decoration either. It is where the report is assembled with its bibliography, where partial results are labelled as partial, and where the errors list becomes something a reader can see:

Python
def finish(state: ResearchState) -> dict:    if not state["draft"]:        return {"draft": ""}        # nothing verified, so no report — never invent one    by_id = {}    for s in state["sources"]:        by_id[Source(**s).id] = s    used = sorted({c["source_id"] for c in state["verified_claims"]})    biblio = "\n".join(f"[{i}] {by_id[i]['title']} — {by_id[i]['url']}"                       for i in used if i in by_id)    caveats = []    if state["errors"]:        caveats.append(f"{len(state['errors'])} step(s) failed; "                       "coverage may be incomplete.")    if not state["approved"]:        caveats.append("Not approved by review; revisions exhausted.")    return {"draft": state["draft"] + "\n\n## Sources\n" + biblio +                     ("\n\n> " + " ".join(caveats) if caveats else "")}

Running it

Python
config = {"configurable": {"thread_id": "eu-payments-2026"}}initial = {"question": "How has EU payment regulation changed since 2023?",           "subquestions": [], "sources": [], "verified_claims": [],           "draft": "", "critique": "", "approved": False,           "revisions": 0, "tokens_used": 0, "errors": []}for chunk in graph.stream(initial, config, stream_mode="updates"):    for node, update in chunk.items():        print(f"{node:>13} -> {list(update)}")
Text
      planner -> ['subquestions', 'tokens_used']       search -> ['sources', 'tokens_used']       (×4, concurrent)     verifier -> ['verified_claims', 'errors', 'tokens_used']  synthesiser -> ['draft', 'revisions', 'tokens_used']       critic -> ['approved', 'critique']         approved=False  synthesiser -> ['draft', 'revisions', 'tokens_used']       critic -> ['approved', 'critique']         approved=True       finish -> ['draft']

A representative run: planner 2,100 tokens; four searchers, no model tokens but 11 pages fetched; verifier 11 sources × ~2,800 tokens = 30,800; synthesiser 24,600; critic 11,900; one revision cycle adds 24,600 + 11,900. Total 2,100 + 30,800 + 24,600 + 11,900 + 24,600 + 11,900 = 105,900 tokens, comfortably inside 200,000.

At example rates of roughly 3 dollars per million input tokens and 15 per million output (check your model's current prices), with about a 70/30 split, that is (0.7 × 105,900 × 3 + 0.3 × 105,900 × 15) / 1,000,000 = 0.22 + 0.48 = 0.70 dollars per report. The first, unguarded run cost 41 dollars — about 59 times as much — and the only difference is three comparisons in a routing function.

Wall-clock: planner 1.9 s, four searchers concurrently at max(3.1, 4.4, 2.8, 6.2) = 6.2 s, verifier 14.0 s, synthesiser 8.1 s, critic 3.3 s, revision cycle 11.4 s, finish 0.1 s — about 45 seconds, inside the 3-minute requirement. Had the searchers run sequentially they would have taken 3.1 + 4.4 + 2.8 + 6.2 = 16.5 s instead of 6.2, adding 10.3 seconds for nothing.

Step 5: defensive error handling

The governing principle: a node should return an error into state rather than raise, unless the error means no useful output is possible. A raised exception in one searcher aborts the whole graph and throws away the three searches that succeeded.

Python
import functools, timedef resilient(node_name: str, fallback: dict):    """Wrap a node so failures become state, not exceptions."""    def decorator(fn):        @functools.wraps(fn)        def wrapper(state):            t0 = time.time()            try:                return fn(state)            except Exception as exc:                return {**fallback,                        "errors": [{"node": node_name,                                    "error": f"{type(exc).__name__}: {exc}",                                    "elapsed_s": round(time.time() - t0, 2)}]}        return wrapper    return decoratorverifier = resilient("verifier", {"verified_claims": []})(verifier)critic   = resilient("critic", {"approved": True, "critique": ""})(critic)

Choosing each fallback is a judgement about which way to fail. The verifier's fallback is an empty claim list, which routes to finish and produces no report — correct, because a report with unverified claims is worse than no report. The critic's fallback is approved: True, which ships the draft unreviewed — correct, because the mechanical citation checks have already run and a broken critic should not block delivery indefinitely.

FailureResponseUser-visible outcome
One searcher throwsRecord error, continue with other resultsReport with a coverage caveat
All searchers return nothingRoute to finish"No sources found", not a fabricated report
Model returns non-JSONCatch, degrade to a safe defaultFewer sub-questions, still a report
Quote not found in sourceDrop the claim, record the errorReport without that unverifiable claim
Critic never approvesRevision cap routes to finishBest available draft, marked unapproved
Budget 85% consumedStop revising, finishReport inside cost limits
Process killed mid-runCheckpoint resumegraph.invoke(None, config) continues

Two more defences worth adding before this touches real users. A per-node timeout, because a hung fetch inside a searcher blocks the superstep and therefore the whole graph — wrap the tool call, not the node, so the searcher can return partial results. And a circuit breaker on the search tool: after five consecutive failures, stop calling it for sixty seconds and let searchers return empty immediately, so that an outage at the search provider produces fast degraded reports rather than a slow queue of timeouts.

Where people get it wrong

Building agents before the state schema. This produced the results/sources mismatch and it produces one of these in nearly every first build. Write the TypedDict first and make every agent's signature reference it.

Trusting model JSON. Every json.loads on model output needs a try/except and a schema assertion. The failure is not rare and it is not graceful.

No cap on fan-out. Sub-question count, sources per search, sources sent to the verifier: all three need explicit numeric limits. A prompt saying "3 to 5" is a suggestion.

Letting the critic be the only stopping condition. An LLM critic that can always find something to improve will always find something to improve. Pair it with a hard revision cap and a budget cap.

Reporting partial success as success. Three of four searchers failing should be visible in the output, not just in the logs. The caveat line in finish costs three lines and prevents a reader treating a thin report as a thorough one.

What this means when you build

The five agents in this system are, individually, unremarkable — each one is a prompt, a tool or two, and some parsing. Nearly all the engineering value sits in four places that are not the agents at all: the state schema with its reducers, the routing functions with their caps, the mechanical verification of quotes, and the assembly step that tells the truth about what failed.

That ratio holds generally. When a multi-agent build goes wrong, the cause is almost never that an agent reasoned badly. It is that two agents disagreed about a key name, or a loop had no counter, or nobody checked something checkable, or a partial result was presented as a complete one.

So build in this order and resist the urge to reorder it: draw the graph and label the edges, write the state schema and choose every reducer deliberately, write the guards, and only then write the agents. The agents are the part you can iterate on cheaply — a prompt change is a minute's work. The state schema is the part that is expensive to change once five nodes depend on it, and the guards are the part whose absence you discover through a 41-dollar invoice.