Multi-Agent Systems and Collaboration

Observability for Multi-Agent Systems


At 09:40 a support lead reported that roughly a fifth of research requests were returning nothing — no answer, no error, just a spinner that eventually gave up. The on-call engineer opened the dashboards. Every agent showed 99.9%+ uptime. Every health check was green. CPU was flat, memory was flat, error rates were near zero across all six services.

The measured failure rate at the user's end was 22%.

Run the arithmetic that the dashboards implied. Six agents, each 99.9% available, all needed for a request to succeed:

0.9996=0.9940⇒expected failure rate=0.60%0.999^6 = 0.9940 \quad\Rightarrow\quad \text{expected failure rate} = 0.60\%

Expected 0.6%, observed 22% — a factor of 37. The dashboards were not lying about the agents. They were measuring the wrong thing entirely. Every one of those failures happened between agents: a handoff whose recipient was at capacity and declined, a queued task whose result nobody collected, a delegation that timed out while the parent had already given up. None of that is visible in per-service uptime, because none of it happens inside a service.

The green dashboards that missed itCoordination healthInter-agentmessage latency, p99Coordination success rateQueue depth and ageTasks handedoff, never returnedRetries percompleted request
Every agent at 99.9 per cent uptime says nothing about whether the handoffs between them completed.

What is genuinely different here

Single-agent observability is well understood: log the prompt, the tool calls, the tokens, the latency, the outcome. All of it happens in one process and one transcript. Multi-agent systems add three failure classes that have no single-agent equivalent.

Failure classExampleWhy standard monitoring misses it
Failures in the gapsHandoff declined; result never collectedNo service was unhealthy during it
Emergent behaviourTwo agents ping-ponging a task 40 timesEvery individual call succeeded quickly
Partial success3 of 4 research agents returned; report written from 3Reported as success; quality silently degraded

The unifying property is that a multi-agent system can be entirely healthy component-wise and entirely broken end to end. So the observability you need is not more per-agent metrics. It is metrics about the relationships.

In a multi-agent system, the interesting failures do not happen inside agents. They happen in the spaces between them, where nothing is being monitored.

Inter-agent message latency

Define it precisely, because a vague definition produces a useless metric. There are three distinct intervals and you want all three separately:

Text
   agent A                queue                    agent B      │                     │                         │      ├── sent_at ─────────►│                         │      │       (1) transit   ├── enqueued_at           │      │                     │   (2) queue wait        │      │                     ├── dequeued_at ─────────►│      │                     │                         ├── started_at      │                     │                         │   (3) processing      │                     │                         ├── completed_at      │◄──────────────── total handoff ───────────────┤

Transit is network and serialisation — usually milliseconds. Queue wait is the backlog. Processing is the agent's own work. Teams commonly measure only (3), which is the one thing that was never the problem.

Python
import timefrom dataclasses import dataclass, field@dataclassclass Envelope:    payload: dict    from_agent: str    to_agent: str    trace_id: str    sent_at: float = field(default_factory=time.time)    enqueued_at: float | None = None      # broker accepted the message    dequeued_at: float | None = None      # a worker for agent B took it    started_at: float | None = None    completed_at: float | None = None    def timings(self) -> dict[str, float]:        return {            "transit_s":    (self.enqueued_at or 0) - self.sent_at,            "queue_wait_s": (self.dequeued_at or 0) - (self.enqueued_at or 0),            "processing_s": (self.completed_at or 0) - (self.started_at or 0),            "total_s":      (self.completed_at or 0) - self.sent_at,        }def record_handoff(env: Envelope, metrics):    t = env.timings()    labels = {"from": env.from_agent, "to": env.to_agent}    for name, value in t.items():        metrics.observe(f"handoff_{name}", value, labels)

Two practical warnings. First, these timestamps come from different machines — sent_at from agent A, dequeued_at from agent B's worker — so any interval that crosses machines includes clock skew and can even go negative. Either accept that it is approximate, or synchronise clocks and treat sub-50-millisecond transit numbers as noise. Second, label by the pair of agents, not just the receiver. "Handoffs into the fraud agent are slow" is much less useful than "handoffs from triage to fraud are slow while handoffs from intake to fraud are fast" — which immediately points at the brief triage is building.

Why percentiles, with numbers

Take 100 handoffs. Ninety-five complete in 0.2 seconds; five take 30 seconds because they queue behind a slow agent.

  • Mean: (95 × 0.2 + 5 × 30) / 100 = (19 + 150) / 100 = 1.69 seconds
  • Median (p50): 0.2 seconds
  • p95: 0.2 seconds — the 95th value is still a fast one
  • p99: 30 seconds

The median says everything is instant. The mean says 1.69 seconds, a number that describes no actual request. Only p99 shows the 30-second stall that 5% of your users experienced. And in a system where a request touches six agents, a 5% chance of a 30-second stall per hop means the probability of at least one stall is 1 - 0.95^6 = 0.265, or 26.5%. A per-hop tail that looks negligible becomes a one-request-in-four problem end to end.

Track p50, p95 and p99 for every agent pair. The mean of a bimodal latency distribution describes a request that never happens.

Coordination success rate: the metric that was missing

"Agent uptime" answers "was the process running?". The question you actually care about is "did the work get through?". These come apart badly, and the 22% incident is what that looks like.

Four rates, each measuring a different gap:

MetricDefinitionDetectsHealthy
Handoff success rateaccepted ÷ offeredCapacity limits, bad briefs, missing capabilities> 95%
Task completion ratecompleted ÷ startedWork that enters and never leaves> 98%
Escalation rateescalated ÷ startedGuards firing: depth caps, deadlines, no-consensus< 3%
Re-delegation ratehops > 1 ÷ totalPing-pong between agents; wrong initial routing< 10%
Python
from collections import defaultdictclass CoordinationMetrics:    def __init__(self):        self.offered   = defaultdict(int)   # (from, to) -> count        self.accepted  = defaultdict(int)        self.declined  = defaultdict(lambda: defaultdict(int))  # pair -> reason        self.completed = defaultdict(int)    def handoff_success_rate(self, pair) -> float | None:        n = self.offered[pair]        return self.accepted[pair] / n if n else None    def worst_pairs(self, min_samples=20, k=5):        rows = [(p, self.handoff_success_rate(p), self.offered[p])                for p in self.offered if self.offered[p] >= min_samples]        rows = [r for r in rows if r[1] is not None]        return sorted(rows, key=lambda r: r[1])[:k]

The declined breakdown by reason is what turns the metric into a fix. In the 22% incident it showed this:

Agent pairOfferedAcceptedRateTop decline reason
planner → search4,1204,08899.2%at capacity
search → verifier3,9603,10278.3%brief missing 'source_urls'
verifier → writer3,1023,09099.6%at capacity

Multiply the chain: 0.992 × 0.783 × 0.996 = 0.774, so 22.6% of requests died somewhere — matching the observed 22% almost exactly. And the cause is named in the table: the search agent had started omitting source_urls from its brief after a change to its output schema. Nothing was down. One field went missing.

This is the general shape. End-to-end success is the product of the per-hop success rates, so a single hop at 78% ruins a chain of otherwise excellent links, and only per-pair measurement finds it.

Per-agent health checks that mean something

A health check returning {"status": "ok"} because the process is running is worse than no health check, because it produces confident green dashboards during an outage. Distinguish three levels:

LevelQuestionUsed byCost
LivenessIs the process alive and not deadlocked?The orchestrator, to restart itMicroseconds
ReadinessCan it accept work right now?The router, to decide where to sendMilliseconds
CapabilityCan it actually do its job end to end?Deploy gates and periodic probesSeconds and tokens
Python
class AgentHealth:    def __init__(self, agent, max_queue=50):        self.agent, self.max_queue = agent, max_queue        self.last_success_ts = time.time()        self.consecutive_failures = 0    def liveness(self) -> dict:        # Only: is the loop turning? No external calls.        stale = time.time() - self.agent.last_loop_tick        return {"alive": stale < 30, "loop_stale_s": round(stale, 1)}    def readiness(self) -> dict:        reasons = []        if self.agent.queue_depth >= self.max_queue:            reasons.append("queue full")        if self.consecutive_failures >= 5:            reasons.append("circuit open")        if time.time() - self.last_success_ts > 600:            reasons.append("no success in 10 minutes")        return {"ready": not reasons, "reasons": reasons,                "queue_depth": self.agent.queue_depth}    def capability(self) -> dict:        # A real, tiny task through the real path. Run every few minutes.        try:            t0 = time.time()            out = self.agent.run(CANARY_TASK, budget_tokens=400)            ok = CANARY_EXPECTED in out.get("answer", "")            return {"capable": ok, "latency_s": round(time.time() - t0, 2)}        except Exception as exc:            return {"capable": False, "error": f"{type(exc).__name__}: {exc}"}

The capability check is the one that catches the interesting failures: an expired API key, a model deprecation, a tool whose upstream schema changed. All three leave liveness and readiness perfectly green. Keep the canary task small and deterministic — a fixed question with a known substring in the answer — and run it on a timer rather than per request, so it costs a few hundred tokens per agent per five minutes rather than doubling your bill.

Notice that readiness returns reasons. A boolean tells you an agent is not taking work; the reason tells you whether to scale it, restart it, or look upstream.

Tracing a request across agents, queues and retries

Metrics tell you something is wrong. Traces tell you where. The requirement is a single identifier that survives every hop, including hops through a queue and hops through a retry — which is exactly where naive tracing breaks, because a queue is a process boundary and most tracing libraries lose context across it.

Python
import uuid, contextvars, time_ctx = contextvars.ContextVar("trace_ctx", default=None)class Span:    def __init__(self, name, kind, parent=None, trace_id=None):        self.trace_id = trace_id or (parent.trace_id if parent                                     else uuid.uuid4().hex)        self.span_id = uuid.uuid4().hex[:16]        self.parent_id = parent.span_id if parent else None        self.name, self.kind = name, kind        self.attrs: dict = {}        self.start = time.time()        self.end: float | None = None    def finish(self, **attrs):        self.attrs.update(attrs)        self.end = time.time()        EXPORTER.export(self)    def carrier(self) -> dict:        """Inject into a queue payload so the consumer can continue the trace."""        return {"trace_id": self.trace_id, "parent_span_id": self.span_id}def enqueue_with_trace(queue, fn, payload, span: Span):    payload = {**payload, "_trace": span.carrier()}     # travels with the task    return queue.enqueue(fn, payload)def run_task_with_trace(payload, agent_name):    carrier = payload.pop("_trace", {})    parent = _FakeParent(**carrier) if carrier else None    span = Span(f"{agent_name}.run", kind="agent", parent=parent)    try:        result = AGENTS[agent_name].run(payload)        span.finish(status="ok", tokens=result["tokens"],                    tool_calls=len(result["tool_calls"]))        return result    except Exception as exc:        span.finish(status="error", error=type(exc).__name__)        raise

The _trace key inside the task payload is the whole trick. Trace context is normally propagated in HTTP headers; a queue has no headers, so you put it in the message body. Without this, every trace ends at the enqueue call and starts fresh in the worker, and you get dozens of one-span traces instead of one useful waterfall.

Here is what the resulting trace showed for a slow request in the incident:

Text
trace 9f2c…  total 41.8s├─ planner.run                  1.9s   tokens=2,140├─ handoff planner→search       0.1s├─ search.run                   3.4s   tokens=6,880  tool_calls=8├─ handoff search→verifier     22.6s   ◄── queue wait, not agent time│   └─ attempt 1 declined  reason="brief missing 'source_urls'"│   └─ attempt 2 declined  reason="brief missing 'source_urls'"│   └─ attempt 3 accepted (fallback brief)├─ verifier.run                 9.2s   tokens=11,400└─ writer.run                   4.6s   tokens=8,020

Total agent processing is 1.9 + 3.4 + 9.2 + 4.6 = 19.1 seconds. Total elapsed is 41.8. So 54% of the request's life was spent in one handoff, retrying against a validation failure. No per-agent dashboard could show that, because no agent was slow.

Attributes worth attaching to every agent span, because they are the ones you will wish you had: agent name and version, model name, prompt and completion tokens, number of tool calls, retry attempt number, delegation depth, and the decline reason where applicable. Tokens on the span are what let you answer "which agent is burning the budget" without a separate cost pipeline.

Alerting on coordination, not infrastructure

Infrastructure alerts — CPU, memory, process down — were all silent during the incident, and they were right to be. The alerts you need fire on relationships.

SymptomLikely causeFirst fix
Handoff success rate for one pair drops below 90%Brief schema changed; receiver at capacityRead the top decline reason for that pair
p99 handoff latency > 10× p50Bimodal queueing behind one slow consumerCheck queue depth and worker count for the receiver
Re-delegation rate risesRouting is choosing the wrong first agentInspect the router's classification distribution
Escalation rate rises with no error rate changeA guard is firing: depth cap, deadline, no consensusBreak escalations down by guard
Tokens per request rise, success rate flatAn agent is looping to its step capLook at tool_calls per span, p99
Completion rate falls while all agents healthyResults produced but never collectedCheck result TTL and the collector's liveness

Two rules make these alerts survivable. Alert on change, not thresholds: "handoff success for this pair fell more than 5 points versus the 7-day baseline" fires on real regressions and stays quiet during normal variation, whereas "below 95%" fires forever on a pair that legitimately sits at 94%. And require a minimum sample size — a pair with 3 handoffs in the window and 1 decline is at 67%, which is noise, and an alert that fires on it will be muted by the second week.

Where people get it wrong

Measuring agents instead of edges. The default instinct is per-service dashboards, because that is what the tooling gives you. Every failure in the opening incident lived on an edge. Add per-pair metrics or you cannot see them.

Averaging latency. A mean over a bimodal distribution describes nothing. It also hides exactly the failures users notice, because users notice the tail.

Shallow health checks. "The process responds to HTTP" is not health. An agent with a revoked API key answers that check instantly and cannot do a single unit of work.

Losing the trace at the queue. Traces that stop at enqueue turn one 42-second story into six unrelated fragments. Put the context in the payload.

Counting partial success as success. A report written from 3 of 4 research agents returns 200 and looks fine. Record the denominator: log how many contributors were expected and how many arrived, and alert when the ratio drops.

Sampling traces uniformly. At 1% sampling you will almost never capture the rare 30-second handoff, which is the one you need. Sample all errors, all slow requests, and 1% of the rest.

What this means when you build

Generate a trace ID at the system's edge — the HTTP handler, the webhook, the scheduler tick — and make it a required field on every envelope, task payload and state object. Not optional, not defaulted: required, so that code which forgets it fails at construction rather than producing an orphan trace you discover during an incident.

Instrument the edges before the nodes. Per-agent latency and token counts are easy and already partly covered by your model provider's dashboards. Per-pair handoff success, decline reasons and queue wait are the numbers nothing gives you for free, and they are the ones that explain a 22% failure rate that six green services cannot.

Build one view that answers "what happened to request X" in a single query. In practice that means one table where every row is one span with trace_id, parent_span_id, agent, phase, duration, status and tokens. It is unglamorous and it is the difference between diagnosing an incident in four minutes and correlating three log streams by eye for two hours.

And write the decline reason down every single time a handoff is refused. It costs one string. In the incident above, that one string — brief missing 'source_urls' — was the entire diagnosis, and without it the team had a 22% failure rate, six healthy services, and nowhere to start.