Multi-Agent Systems and Collaboration

Evaluating Emergent Behaviors and Performance


A research-assistant system ran in production for three weeks without a single deployment. Over those weeks the median tokens per report climbed from 106,000 to 189,000 — a 78% rise — while the human quality rating stayed flat at 3.8 out of 5. The bill nearly doubled for identical output.

No code changed. No prompt changed. No configuration changed. Every agent behaved exactly as written.

The state logs showed the mechanism. A page cache had been added to the fetch tool early on, so repeat fetches were free. As the cache filled, searchers began returning more overlapping sources — cheap ones they had seen before ranked well. The verifier then produced more near-duplicate claims from those overlapping sources. The critic, reading a draft built from near-duplicate claims, flagged it as repetitive and sent it back. Average revisions per report rose from 0.4 to 1.7, and each revision costs about 36,000 tokens.

0.4 → 1.7 revisions is 1.3 × 36,000 ≈ 47,000 extra tokens, and the rest came from the verifier working through a longer source list. Every individual component was correct. The behaviour belonged to the system, not to any agent in it.

Tokens per report over three weeks, in thousands106124141158172189012345week 1 median78 percent moreNo deployment happened in this window, and the human quality rating stayed flat at 3.8 out of 5.
Cost drift is emergent: agents pass each other steadily longer context until the bill doubles for identical output.

What counts as emergent behaviour

Emergent behaviour is a system-level pattern that no agent was programmed to produce, arising from the interaction of agents each following its own local rules. The test is simple: can you point at the line of code that causes it? If yes, it is a bug or a feature. If no — if it only exists in the interaction — it is emergent.

BugDesigned behaviourEmergent behaviour
Located inOne functionOne functionThe interaction between several
Reproducible in isolationYesYesNo — needs the whole system
Fixed byEditing that functionn/aChanging the interaction rules
Found byUnit testsSpecificationAnalysing run logs in aggregate

Emergence is not automatically bad. Three categories are worth separating:

Beneficial. Spontaneous specialisation: with capability-based routing and a skill score that updates from outcomes, agents drift toward the task types they do well, without anyone assigning roles. Self-balancing load: least-loaded routing plus varying task sizes produces even utilisation that no scheduler explicitly computed.

Harmful. Oscillation: two agents hand a task back and forth, each locally correct. Convergence collapse: agents that read each other's outputs stop disagreeing, so a five-agent vote carries the information of one. Cost spirals: the story above. Cascading retries: one slow agent causes timeouts upstream, which cause retries, which increase load, which slows it further.

Neutral but worth knowing. Ordering preferences that stabilise — one searcher always finishing first, so its sources always appear at the top of the verifier's list and get checked before the budget runs out. Harmless until the day that agent is slow.

If you cannot name the function that causes a behaviour, do not look for a bug. Look at the interaction, and look at it in aggregate over many runs.

Detecting emergence from a state log

Detection is only possible if you logged the right things. The minimum is one row per state transition with enough context to reconstruct the run:

Python
from dataclasses import dataclass, asdictimport json, time@dataclassclass Transition:    run_id: str    seq: int                 # monotonic within the run    node: str                # which agent produced this    keys_written: list[str]    tokens: int    elapsed_s: float    ts: float    # cheap derived signals, computed at write time    sources_total: int = 0    sources_unique: int = 0    revisions: int = 0    errors: int = 0def log_transition(sink, run_id, seq, node, update, state, elapsed):    urls = [s["url"] for s in state.get("sources", [])]    sink.write(json.dumps(asdict(Transition(        run_id=run_id, seq=seq, node=node,        keys_written=sorted(update.keys()),        tokens=update.get("tokens_used", 0), elapsed_s=elapsed, ts=time.time(),        sources_total=len(urls), sources_unique=len(set(urls)),        revisions=state.get("revisions", 0),        errors=len(state.get("errors", []))))) + "\n")

With that, four detectors cover most of what goes wrong.

1. Cycle and oscillation detection

Python
from collections import Counterdef detect_oscillation(nodes: list[str], min_repeats: int = 3):    """Find a short node sequence that repeats back to back."""    findings = []    for window in (2, 3, 4):        for i in range(len(nodes) - window * min_repeats + 1):            pattern = tuple(nodes[i:i + window])            repeats = 1            j = i + window            while tuple(nodes[j:j + window]) == pattern:                repeats += 1                j += window            if repeats >= min_repeats:                findings.append({"pattern": pattern, "repeats": repeats,                                 "start_seq": i})    return findingsdetect_oscillation(["planner", "search", "verifier",                    "synthesiser", "critic", "synthesiser", "critic",                    "synthesiser", "critic"])# [{'pattern': ('synthesiser', 'critic'), 'repeats': 3, 'start_seq': 3}]

This is what a runaway revision loop looks like in data. Run it across all runs and count how many contain a pattern with three or more repeats; a rising fraction is the earliest signal of a cost spiral.

2. Diversity collapse

The direct measure for the cache problem: what fraction of gathered sources are distinct?

Python
def source_diversity(transitions) -> float:    last = max(transitions, key=lambda t: t["seq"])    return (last["sources_unique"] / last["sources_total"]            if last["sources_total"] else 1.0)

In week one this sat at 0.90 — 20 sources gathered, 18 distinct. By week three it was 0.58: 31 gathered, 18 distinct. The same 18 useful pages, found repeatedly, each one verified again. That single ratio explains most of the token growth and is one line to compute.

3. Cost drift

Compare a recent window against a baseline, using medians so one runaway run does not move the number:

Python
import statistics as statsdef cost_drift(run_tokens: list[int], baseline_n=200, recent_n=50):    if len(run_tokens) < baseline_n + recent_n:        return None    base = stats.median(run_tokens[:baseline_n])    recent = stats.median(run_tokens[-recent_n:])    return {"baseline": base, "recent": recent,            "drift_pct": round((recent - base) / base * 100, 1)}# {'baseline': 106000, 'recent': 189000, 'drift_pct': 78.3}

A drift above about 15% with flat quality is the alert. The three weeks of gradual growth were invisible day to day — roughly 3% per day is under the noise floor — and unmistakable over a fortnight.

4. Diversity of agent opinion

Where several agents vote or review, measure how often they disagree. Falling disagreement is not consensus improving; it is usually agents anchoring on shared context.

Python
def disagreement_rate(votes_per_run: list[list[str]]) -> float:    """Fraction of runs where not all voters agreed."""    return sum(len(set(v)) > 1 for v in votes_per_run) / len(votes_per_run)

If five reviewers disagree on 40% of runs in week one and 6% in week five, your five-agent panel now has roughly the information content of one agent, and you are paying five times over for it.

Wiring the detectors in

Two placements, doing different jobs.

Online, as a guard. Cheap checks inside the routing function, stopping a spiral in progress:

Python
def route_after_critic(state) -> str:    if state["approved"] or state["revisions"] >= MAX_REVISIONS:        return "finish"    if state["tokens_used"] > TOKEN_BUDGET * 0.85:        return "finish"    # Emergence guard: if the draft barely changed, revising again is futile.    if similarity(state["draft"], state.get("previous_draft", "")) > 0.95:        return "finish"    return "synthesiser"

Offline, as a nightly job. The expensive aggregate analysis — oscillation frequency, diversity trend, cost drift, disagreement rate — over the last few hundred runs, written to a table you look at weekly. The cache-induced spiral could only ever be caught here, because no single run looked wrong.

Performance metrics

Five numbers, and the third one is the one most teams compute incorrectly.

MetricDefinitionReported as
LatencyRequest received → result deliveredp50, p95, p99 — never the mean
ThroughputCompleted tasks per unit time at a given concurrencyTasks/minute, with the concurrency stated
Cost per successful taskTotal spend ÷ successful tasksCurrency per success
Success rateTasks meeting the acceptance criteria ÷ tasks startedPercentage, with the criteria written down
QualityScore against a rubric, on a fixed evaluation setMean and distribution

Cost per successful task is the honest one. Take 1,000 runs at 0.70 dollars each: total spend 700 dollars. If 88% succeed, that is 880 successes, so the real cost is 700 / 880 = 0.795 dollars per useful result — 13.6% higher than the naive figure. Now suppose an optimisation cuts per-run cost to 0.60 but drops success to 76%: 600 / 760 = 0.789 dollars. Almost no improvement, and you have made a quarter of your users' requests fail. Only the corrected metric shows that.

Divide by successes, not by attempts. Failed runs consume tokens and deliver nothing, so a change that trades success rate for per-run cost usually saves nothing at all.

On quality: fix an evaluation set of 30 to 50 questions and score every candidate configuration against the same set. Rubric items for a research system might be — does every factual sentence carry a resolvable citation, do the quotes appear in the cited sources, is the question actually answered, is anything asserted that no claim supports. Three of those four are mechanical, which means most of your quality score can be computed rather than rated.

Coordination metrics

Performance metrics describe outcomes. Coordination metrics describe whether the multi-agent structure is earning its keep.

Parallel efficiency

Four searchers take 3.1, 4.4, 2.8 and 6.2 seconds. Sequentially that is 16.5 seconds; run concurrently the stage takes max = 6.2.

S=16.56.2=2.66E=Sn=2.664=0.665S = \frac{16.5}{6.2} = 2.66 \qquad E = \frac{S}{n} = \frac{2.66}{4} = 0.665

Speed-up 2.66×, efficiency 66.5%. The missing third is the straggler: three agents finish and wait for the fourth. Efficiency below about 50% means adding parallel agents is mostly buying idle time, and the fix is either to even out the work or to proceed with partial results when most agents have finished.

Coordination overhead

What fraction of wall-clock time is not spent doing model or tool work? For the 45-second run measured earlier, model and tool time totalled 38.4 seconds:

overhead=45.0−38.445.0=14.7%\text{overhead} = \frac{45.0 - 38.4}{45.0} = 14.7\%

Under 15% is healthy for a five-agent workflow. Above 30% means the structure costs more than it returns, and the honest response is to merge agents.

Redundancy rate

Work done more than once: duplicate fetches, duplicate verifications, near-duplicate claims. This is 1 - diversity from the detector above, and it is the metric that would have caught the cache spiral in week one.

Coordination metricHealthyWhat a bad value means
Parallel efficiency> 0.6Stragglers, or uneven task sizes
Coordination overhead< 15%Too many agents for the work
Redundancy rate< 15%Agents duplicating each other's work
Handoff success rate> 95%Brief mismatch, or receivers at capacity
Agent idle fraction< 30%Over-provisioned, or a bottleneck upstream

Comparing centralised and decentralised coordination on this system

The abstract trade-off — a coordinator gives control and becomes a bottleneck; peers give resilience and duplicate work — is settled by measurement on your workload, not by argument. Here is how to run that comparison properly.

Design. Same 40-question evaluation set, 200 runs per arm, interleaved rather than run in separate blocks so that a slow afternoon at the search provider hits both arms equally. Identical models, identical prompts, identical caps. The only difference is who decides which searcher takes which sub-question: a coordinator node, or searchers claiming sub-questions from a shared Redis key with SET NX.

MetricCentralisedDecentralisedReading
p50 latency41 s38 sPeers are slightly faster in the typical case
p95 latency78 s142 sPeers have a much worse tail
Tokens per report106,000131,000+23.6% from duplicated work
Success rate94%89%Claim races leave some sub-questions unanswered
Redundancy rate3%19%The direct cause of the token gap
Cost per success0.740.98+32%
Behaviour when one agent diesRun stalls until reclaimClaim expires, peer continuesPeers win on resilience

Two things stand out. The p50 comparison favours peers and the p95 comparison reverses it decisively — a reminder that reporting a single latency number would have led to the wrong conclusion. And the token gap is fully explained by the redundancy rate: 19% duplicated work against 3% predicts roughly a 16% overhead, and the observed gap of 23.6% is that plus the extra verification those duplicates trigger.

Turning it into a recommendation

Weight the metrics by what the product actually cares about, normalise so that 1.0 is the better arm, and score. For a system where cost matters most, then tail latency, then success:

MetricWeightCentralised (normalised)Decentralised (normalised)
Cost per success0.41.0000.74 / 0.98 = 0.755
p95 latency0.31.00078 / 142 = 0.549
Success rate0.31.0000.89 / 0.94 = 0.947
Weighted total1.01.0000.751

Decentralised scores 0.4(0.755) + 0.3(0.549) + 0.3(0.947) = 0.302 + 0.165 + 0.284 = 0.751. Centralised wins clearly at this scale, and the recommendation should say so with its conditions attached: at four searchers, with a coordinator decision time of about 0.12 seconds and a searcher duration of about 4 seconds, the coordinator is nowhere near its ceiling of 4 / 0.12 ≈ 33 workers. Above roughly 33 searchers the coordinator saturates and the comparison flips. So the honest recommendation is "centralised now; revisit if fan-out exceeds about 25".

A hybrid is usually the answer that neither arm represents: a coordinator that owns allocation and the token budget, with peers free to talk directly for everything else. That keeps redundancy near 3% while removing the coordinator from the critical path of every interaction.

Where people get it wrong

Averaging latency. The mean of a bimodal distribution describes a request nobody makes. Report percentiles; the p50/p95 reversal above is exactly why.

Cost per attempt instead of cost per success. Makes any change that trades reliability for token count look like a win.

Evaluating on whatever ran that day. Traffic mix changes, so week-on-week comparisons on production traffic measure the traffic, not the system. Use a fixed evaluation set.

Reading emergence one run at a time. The three-week spiral was invisible in every individual run and obvious in the aggregate. Emergent behaviour is a property of distributions.

Calling any surprise "emergence". Most surprises are bugs. Apply the test: can you name the function? A revision loop with no cap is a missing guard, not an emergent phenomenon, and calling it emergence discourages you from fixing it properly.

Assuming more agents means better quality. Only when their errors are independent. Measure disagreement rate; when it collapses, the extra agents are adding cost and no information.

Running A/B arms sequentially. Arm A in the morning and arm B in the afternoon measures the provider's afternoon latency as much as your architecture. Interleave.

What this means when you build

Log state transitions from the first day, with the derived signals computed at write time — source counts, revision counts, tokens, error counts. Every detector in this lesson is a few lines over that log, and none of them can be run retrospectively against logs that were never written. The three-week spiral was diagnosable in an afternoon because those rows existed; without them it would have been a guess.

Put four numbers on a weekly review, not a real-time dashboard: median tokens per successful report, source diversity, oscillation frequency, and disagreement rate. All four drift slowly, which means real-time alerting will never fire on them and a weekly glance will catch every one. Slow drift is the characteristic signature of emergent behaviour, and it is the one thing conventional monitoring is structurally unable to see.

Make every architectural argument an experiment. Centralised versus decentralised, four searchers versus eight, one critic versus a panel of three — each of these is 200 runs against a fixed evaluation set and half a day. The results are frequently counter-intuitive, and the alternative is choosing your architecture from a diagram.

Finally, attach conditions to every recommendation. "Centralised is better" is false in general and true at four searchers with a 0.12-second decision time. Writing down the conditions and the crossover point turns a conclusion that expires silently into one that tells you when to re-examine it — which matters, because the system that grew 78% more expensive over three weeks did so without anybody changing a line.