Multi-Agent Systems and Collaboration

Single vs Multi-Agent Paradigms


A fintech support team built an assistant. It started with four tools: look up an account, read the transaction log, check card status, open a ticket. It worked beautifully. Six months later it had forty-seven tools — refund processing, dispute filing, KYC document checks, fraud scoring, address changes, statement generation, chargeback rules, and thirty more. The system prompt had grown to eleven thousand tokens because every tool needed usage notes and every edge case needed a warning.

And it had quietly become useless. Customers asking to freeze a stolen card got the fraud-scoring tool instead of the freeze tool. Refund requests triggered dispute filings. The engineers added more instructions — "IMPORTANT: freeze_card is for lost/stolen, NOT fraud_score" — and each new instruction made some other case worse. The failure was not the model being stupid. The failure was that one agent was being asked to hold forty-seven jobs in its head at once.

The fix was not a better prompt. The fix was structural: split the assistant into a router and four specialists — cards, payments, compliance, and account admin — each with about a dozen tools and a prompt that only described its own domain. That is the moment a system stops being single-agent and becomes multi-agent, and it is worth understanding precisely what you gain and what you pay.

Forty-seven tools in one loop, or splitOne agent, one loop• Shared context, no handoff to lose• Tool descriptions crowd the prompt• Selection accuracy falls as tools grow• One failure ends the whole runSeveral specialised agents• Each prompt lists only its own tools• Work runs in parallel where it can• Coordination is now code you own• Context must be passed on explicitly
Splitting buys back selection accuracy and parallelism, and charges you in messages, latency and lost context.

What a single agent actually is

An agent, in this context, is a language model wrapped in a loop that lets it take actions. Strip away the marketing and the loop is about ten lines:

Python
def run_agent(user_message, tools, system_prompt, max_steps=15):    messages = [{"role": "system", "content": system_prompt},                {"role": "user", "content": user_message}]    for step in range(max_steps):        reply = model.generate(messages, tools=tools)        messages.append(reply)        if not reply.tool_calls:            return reply.content            # the agent decided it is done        for call in reply.tool_calls:            result = tools[call.name](**call.arguments)            messages.append({"role": "tool", "tool_call_id": call.id,                             "content": result})    raise StepLimitExceeded("agent did not converge")

That is a single-agent system: one decision-maker, one conversation history, one set of tools, one loop. Everything the system knows at any moment is in that single messages list.

The properties that follow from having one loop

  • One state, no synchronisation. There is exactly one version of the truth. Nothing can go stale, because nothing is copied.
  • Deterministic control flow. Step 4 always happens after step 3. There is no race, no ordering ambiguity, no message arriving out of sequence.
  • Trivially debuggable. The entire behaviour of the system is one transcript you can read top to bottom.
  • Strictly sequential. If the task involves four independent web searches, the agent does them one after another. Four searches at 2.1 seconds each is 8.4 seconds of wall-clock time that could have been 2.1.
  • One point of failure. If the model mis-picks a tool at step 3, every later step is built on that mistake.

Where the single agent breaks — with numbers

Two things degrade as you add tools. Both are measurable, and both bite harder than people expect.

Cost. Tool schemas are re-sent on every model call, because the model is stateless between calls. Suppose our forty-seven tools average 132 tokens of schema each — a name, a description, a few typed parameters with their own descriptions. That is 47 × 132 = 6,204 tokens of schema. A moderately involved support task takes about twelve model calls. So the schemas alone cost 12 × 6,204 = 74,448 tokens, before a single word of the customer's actual problem.

Now split into a router (4 handoff tools, ~530 tokens of schema) and specialists (about 12 tools each, ~1,584 tokens). A typical task is two router calls and ten specialist calls: 2 × 530 + 10 × 1,584 = 1,060 + 15,840 = 16,900 tokens. That is a 77% reduction in schema overhead. You add back roughly 900 tokens per handoff to carry the task brief across, and with three handoffs that is 2,700 tokens — still leaving you far ahead.

Accuracy. This one is worse. Tool-selection accuracy falls as the candidate set grows, and errors compound multiplicatively across a multi-step task. Say the model picks correctly 82% of the time among 47 similar-sounding tools, and 96% of the time among 12 clearly distinct ones. Over a twelve-step task:

0.8212=0.092versus0.9612=0.6130.82^{12} = 0.092 \qquad\text{versus}\qquad 0.96^{12} = 0.613

A 9.2% chance of getting every step right, versus 61.3%. The per-step gap looks like fourteen percentage points. The end-to-end gap is a factor of six and a half. This is the single most under-appreciated number in agent design.

Tool-selection errors compound. A modest drop in per-step accuracy becomes a catastrophic drop in task success, because the task only succeeds if every step succeeds.

What changes when you add a second agent

A multi-agent system is two or more agents, each with its own loop, its own prompt, and its own tools, that exchange information to accomplish something none of them handles alone. The defining feature is not "more models" — it is separate decision loops with separate state.

Here is the support system after the split, in the shape the code actually takes:

Python
router = Agent(    name="router",    system_prompt="Classify the customer request into exactly one domain "                  "and hand off. Do not attempt to solve it yourself.",    tools=[to_cards, to_payments, to_compliance, to_account])cards = Agent(    name="cards",    system_prompt="You handle physical and virtual card issues only: "                  "freeze, replace, activate, set limits, report stolen.",    tools=[freeze_card, replace_card, activate_card, set_limit, ...])# ... payments, compliance, account defined the same waydef handle(request):    domain = router.run(request)          # returns e.g. "cards"    brief  = {"request": request, "customer_id": request.customer_id}    return AGENTS[domain].run(brief)

Notice what the router's prompt does not contain: any description of how to freeze a card. That knowledge lives in one place, and the router cannot get it wrong because the router never sees it.

What you gain

  • Bounded context per agent. Each agent's prompt describes one domain, so it can be specific rather than hedging across forty-seven cases.
  • Genuine parallelism. Independent subtasks run concurrently. Four 2.1-second searches finish in 2.1 seconds instead of 8.4.
  • Isolated failure. If the compliance agent crashes, card freezes keep working. In the single-agent design, one bad tool call poisoned the whole conversation.
  • Independent evolution. The payments team ships a new refund flow without touching the compliance prompt, and without re-testing card freezes.
  • Heterogeneous models. The router is a classification job — a small fast model does it for a fraction of the cost. The compliance agent, where a mistake is expensive, gets the strongest model available.

What you pay

These costs are real and most teams discover them the hard way.

  • Context loss at every boundary. When the router hands to cards, the cards agent does not automatically know that the customer already said "I'm travelling and my flight leaves in an hour". Whatever you forget to put in the brief is simply gone.
  • Coordination overhead. Handoffs cost tokens, latency, and code. A two-agent system with one handoff is fine. An eight-agent system where every agent may call every other has 56 possible edges to reason about.
  • Non-determinism. With parallel agents, two runs of the same input can interleave differently and produce different results. Reproducing a bug becomes a real task.
  • Distributed debugging. There is no single transcript. You need correlation IDs and traces just to reconstruct what happened.
  • Deadlock and livelock. Agent A waits for B, B waits for A. Or two agents hand a task back and forth forever, each convinced it belongs to the other.

Every agent boundary is a place where context can be dropped. Multi-agent design is largely the discipline of deciding what crosses each boundary.

Agent topologies

Once you have several agents, the question is who talks to whom. Three shapes cover almost everything built in practice.

Centralised

One coordinator holds the plan. Workers only talk to the coordinator, never to each other.

Text
              ┌─────────────┐              │ Coordinator │              └──┬───┬───┬──┘        ┌────────┘   │   └────────┐   ┌────▼───┐   ┌────▼───┐   ┌────▼───┐   │Worker A│   │Worker B│   │Worker C│   └────────┘   └────────┘   └────────┘

With n workers this is n communication edges. The coordinator sees everything, so global decisions ("we have spent enough, stop") are easy. It is also the bottleneck: every message queues behind it, and if it dies the system dies.

Decentralised (peer-to-peer)

Agents talk directly to whichever peer they need.

Text
   ┌────────┐◄──────►┌────────┐   │Agent A │        │Agent B │   └───┬────┘        └────┬───┘       │   ┌────────┐     │       └──►│Agent C │◄────┘           └────────┘

Fully connected, n agents give n(n-1)/2 edges. For 3 agents that is 3 — pleasant. For 8 agents it is 28, and for 15 it is 105. Nobody holds a global view, so "have we finished?" becomes a genuinely hard question requiring an explicit termination protocol. In exchange, there is no single point of failure and no central queue.

Hybrid

Clusters coordinate internally around a local lead; leads coordinate with each other. This is how large real systems are built, and how human organisations work: a team has a lead, leads meet, individual contributors mostly talk within their team.

TopologyEdges (n agents)Global view?Single point of failureBest for
CentralisednYes, at coordinatorYes — the coordinatorClear task decomposition, need for global budget or policy control
Decentralisedup to n(n-1)/2NoNoIndependent agents, high fault tolerance, no natural planner
Hybridn + cluster leadsPartial, per clusterPer-cluster onlyLarge systems (10+ agents) with natural groupings

Autonomy is a dial, not a switch

"Autonomous agent" is used as if it were binary. It is not. Autonomy is how much an agent may decide without asking, and you set it per agent, deliberately.

LevelThe agent may…ExampleRight when
ReactiveOnly respond to explicit instructionsA summariser: given text, returns a summaryThe action is cheap, reversible, and fully specified
DeliberativePlan its own multi-step route to a stated goalA research agent choosing which sources to readThe goal is clear but the path is not
ProactiveInitiate work nobody asked forA monitor that opens a ticket when error rate spikesTimeliness matters more than confirmation
GatedPropose, but a human or supervisor approvesA refund agent that drafts refunds over 500 dollars for reviewActions are irreversible or expensive

The common design mistake is giving every agent the same autonomy level. In the fintech example, the cards agent freezing a card is reversible and urgent — high autonomy is correct. The payments agent issuing a refund moves real money — it should be gated above a threshold. Same system, deliberately different dials.

Coordination mechanisms, briefly

Whatever the topology, agents need some mechanism to align. Four appear constantly:

  • Direct messaging. Agent A sends a message addressed to agent B and waits for a reply. Simple, synchronous, and exactly what a function call feels like.
  • Shared state. All agents read and write one structured store — a plan document, a findings dictionary. Nobody addresses anybody; you write what you learned and read what others learned.
  • Publish–subscribe. An agent announces an event ("document_indexed") without knowing who cares. Interested agents subscribe. Publishers and subscribers do not know about each other, which makes adding a fifth agent free.
  • Bidding. The coordinator announces a task, capable agents bid with a cost or confidence, and the best bid wins. Useful when which agent should do a job depends on runtime conditions like current load.

Shared state is the quiet default in most modern frameworks, and it has a specific hazard: two agents writing the same key concurrently. If the search agent and the summary agent both write state["notes"], one write silently disappears. The fix is to make each agent own its own keys, or to define an explicit merge rule (append to a list rather than replacing a value).

A decision procedure you can actually apply

Before splitting anything, answer these in order. The first "no" tells you to stay single-agent.

QuestionIf yesIf no
Does the task contain genuinely independent subtasks?Parallelism will pay for itselfYou get overhead with no speed-up
Do subtasks need different tools, prompts, or models?Specialisation will raise accuracyOne agent with those tools is simpler
Is one agent's prompt now over ~4,000 tokens of instructions?Split by domainPrune the prompt first, do not split
Do different parts need different failure or approval policies?Isolation is worth the costUniform policy is fine in one loop
Do you have tracing across process boundaries?You can debug what you buildBuild tracing before you split

Concretely: a document Q&A bot that retrieves and answers is one agent, and making it three (retriever, reader, writer) adds latency and context loss for nothing. A system that must simultaneously scrape ten sites, cross-check claims, and draft a report with citations is genuinely multi-agent, because the scrapes are independent and the drafter needs a different prompt and a different model than the scrapers.

Do not split an agent because the architecture diagram looks more impressive. Split it when a specific number — token cost, tool-selection accuracy, wall-clock latency — improves.

Named failure modes

These are the four ways teams get this wrong, in rough order of frequency.

Premature decomposition. Three agents on day one for a task one agent handles. Symptom: every agent's prompt is nearly identical, and handoffs pass the entire conversation verbatim. If the brief you pass between agents is "everything", you have built a single agent with extra latency. Merge them.

The thin-brief handoff. The opposite error. The router passes only {"intent": "freeze_card"} and drops the customer's stated urgency, their previous three failed attempts, and the fact that they already tried the app. The specialist re-asks questions the customer already answered. Symptom: users complain the bot "forgot". Fix: define the brief schema explicitly and test it, rather than letting it emerge from whatever the router happened to output.

The chatty mesh. Eight peer agents, each free to message any other. Message volume grows quadratically, latency becomes unpredictable, and no agent knows when the task is done. Symptom: runs that terminate only on the step limit. Fix: impose hierarchy — introduce cluster leads and forbid cross-cluster chatter.

Unbounded loops between peers. The payments agent decides a case is really compliance; compliance decides it is really payments. Each handoff is individually reasonable. Together they never terminate. Fix: carry a hop count in the brief, cap it (three is usually right), and route to a human on overflow.

What this means when you build

Start single-agent. Always. Instrument it before you need to: log every tool call with its arguments, the step index, and whether the result was used in the final answer. Then let the data tell you where to cut.

When you do split, the boundary should follow the tools, not the workflow steps. Splitting "planner → executor → reviewer" is usually a mistake, because all three need the same context and you pay three handoffs to move it around. Splitting "cards / payments / compliance" works because each half owns a disjoint set of tools and a disjoint set of facts, so the brief that crosses the boundary is small.

Write the brief schema before you write the agents. Literally define the dataclass:

Python
from dataclasses import dataclass, field, replace@dataclassclass Handoff:    request: str                    # the customer's words, verbatim    customer_id: str    facts_established: dict         # what the router already confirmed    constraints: list = field(default_factory=list)   # "flight in 1 hour"    hops: int = 0                   # incremented at every handoff    def next_hop(self, **updates) -> "Handoff":        if self.hops >= 3:            raise HandoffLimitExceeded(self.request)        return replace(self, hops=self.hops + 1, **updates)

That single dataclass prevents the thin-brief failure (you can see what you forgot), the unbounded-loop failure (the hop cap is enforced in one place), and it makes handoffs testable without running a model. Roughly half the bugs in early multi-agent systems are bugs in what crosses the boundary, and this is where you fix them.

Finally, pick your topology from the number of agents, not from taste. Two or three agents: centralised, with the coordinator holding the plan. Ten or more with natural groupings: hybrid. Fully peer-to-peer is the right answer far less often than it is chosen, and it is the hardest to debug — reach for it only when you genuinely cannot tolerate a coordinator as a single point of failure.