Course Content
Multi-Agent Systems and Collaboration
4 sections · 12 lessons
Observability for Multi-Agent Systems
At 09:40 a support lead reported that roughly a fifth of research requests were returning nothing — no answer, no error, just a spinner that eventually gave up. The on-call engineer opened the dashboards. Every agent showed 99.9%+ uptime. Every health check was green. CPU was flat, memory was flat, error rates were near zero across all six services.
The measured failure rate at the user's end was 22%.
Run the arithmetic that the dashboards implied. Six agents, each 99.9% available, all needed for a request to succeed:
Expected 0.6%, observed 22% — a factor of 37. The dashboards were not lying about the agents. They were measuring the wrong thing entirely. Every one of those failures happened between agents: a handoff whose recipient was at capacity and declined, a queued task whose result nobody collected, a delegation that timed out while the parent had already given up. None of that is visible in per-service uptime, because none of it happens inside a service.
What is genuinely different here
Single-agent observability is well understood: log the prompt, the tool calls, the tokens, the latency, the outcome. All of it happens in one process and one transcript. Multi-agent systems add three failure classes that have no single-agent equivalent.
| Failure class | Example | Why standard monitoring misses it |
|---|---|---|
| Failures in the gaps | Handoff declined; result never collected | No service was unhealthy during it |
| Emergent behaviour | Two agents ping-ponging a task 40 times | Every individual call succeeded quickly |
| Partial success | 3 of 4 research agents returned; report written from 3 | Reported as success; quality silently degraded |
The unifying property is that a multi-agent system can be entirely healthy component-wise and entirely broken end to end. So the observability you need is not more per-agent metrics. It is metrics about the relationships.
In a multi-agent system, the interesting failures do not happen inside agents. They happen in the spaces between them, where nothing is being monitored.
Inter-agent message latency
Define it precisely, because a vague definition produces a useless metric. There are three distinct intervals and you want all three separately:
agent A queue agent B │ │ │ ├── sent_at ─────────►│ │ │ (1) transit ├── enqueued_at │ │ │ (2) queue wait │ │ ├── dequeued_at ─────────►│ │ │ ├── started_at │ │ │ (3) processing │ │ ├── completed_at │◄──────────────── total handoff ───────────────┤Transit is network and serialisation — usually milliseconds. Queue wait is the backlog. Processing is the agent's own work. Teams commonly measure only (3), which is the one thing that was never the problem.
1import time2from dataclasses import dataclass, field34@dataclass5class Envelope:6 payload: dict7 from_agent: str8 to_agent: str9 trace_id: str10 sent_at: float = field(default_factory=time.time)11 enqueued_at: float | None = None # broker accepted the message12 dequeued_at: float | None = None # a worker for agent B took it13 started_at: float | None = None14 completed_at: float | None = None1516 def timings(self) -> dict[str, float]:17 return {18 "transit_s": (self.enqueued_at or 0) - self.sent_at,19 "queue_wait_s": (self.dequeued_at or 0) - (self.enqueued_at or 0),20 "processing_s": (self.completed_at or 0) - (self.started_at or 0),21 "total_s": (self.completed_at or 0) - self.sent_at,22 }2324def record_handoff(env: Envelope, metrics):25 t = env.timings()26 labels = {"from": env.from_agent, "to": env.to_agent}27 for name, value in t.items():28 metrics.observe(f"handoff_{name}", value, labels)Two practical warnings. First, these timestamps come from different machines — sent_at from agent A, dequeued_at from agent B's worker — so any interval that crosses machines includes clock skew and can even go negative. Either accept that it is approximate, or synchronise clocks and treat sub-50-millisecond transit numbers as noise. Second, label by the pair of agents, not just the receiver. "Handoffs into the fraud agent are slow" is much less useful than "handoffs from triage to fraud are slow while handoffs from intake to fraud are fast" — which immediately points at the brief triage is building.
Why percentiles, with numbers
Take 100 handoffs. Ninety-five complete in 0.2 seconds; five take 30 seconds because they queue behind a slow agent.
- Mean:
(95 × 0.2 + 5 × 30) / 100 = (19 + 150) / 100 = 1.69seconds - Median (p50): 0.2 seconds
- p95: 0.2 seconds — the 95th value is still a fast one
- p99: 30 seconds
The median says everything is instant. The mean says 1.69 seconds, a number that describes no actual request. Only p99 shows the 30-second stall that 5% of your users experienced. And in a system where a request touches six agents, a 5% chance of a 30-second stall per hop means the probability of at least one stall is 1 - 0.95^6 = 0.265, or 26.5%. A per-hop tail that looks negligible becomes a one-request-in-four problem end to end.
Track p50, p95 and p99 for every agent pair. The mean of a bimodal latency distribution describes a request that never happens.
Coordination success rate: the metric that was missing
"Agent uptime" answers "was the process running?". The question you actually care about is "did the work get through?". These come apart badly, and the 22% incident is what that looks like.
Four rates, each measuring a different gap:
| Metric | Definition | Detects | Healthy |
|---|---|---|---|
| Handoff success rate | accepted ÷ offered | Capacity limits, bad briefs, missing capabilities | > 95% |
| Task completion rate | completed ÷ started | Work that enters and never leaves | > 98% |
| Escalation rate | escalated ÷ started | Guards firing: depth caps, deadlines, no-consensus | < 3% |
| Re-delegation rate | hops > 1 ÷ total | Ping-pong between agents; wrong initial routing | < 10% |
1from collections import defaultdict23class CoordinationMetrics:4 def __init__(self):5 self.offered = defaultdict(int) # (from, to) -> count6 self.accepted = defaultdict(int)7 self.declined = defaultdict(lambda: defaultdict(int)) # pair -> reason8 self.completed = defaultdict(int)910 def handoff_success_rate(self, pair) -> float | None:11 n = self.offered[pair]12 return self.accepted[pair] / n if n else None1314 def worst_pairs(self, min_samples=20, k=5):15 rows = [(p, self.handoff_success_rate(p), self.offered[p])16 for p in self.offered if self.offered[p] >= min_samples]17 rows = [r for r in rows if r[1] is not None]18 return sorted(rows, key=lambda r: r[1])[:k]The declined breakdown by reason is what turns the metric into a fix. In the 22% incident it showed this:
| Agent pair | Offered | Accepted | Rate | Top decline reason |
|---|---|---|---|---|
| planner → search | 4,120 | 4,088 | 99.2% | at capacity |
| search → verifier | 3,960 | 3,102 | 78.3% | brief missing 'source_urls' |
| verifier → writer | 3,102 | 3,090 | 99.6% | at capacity |
Multiply the chain: 0.992 × 0.783 × 0.996 = 0.774, so 22.6% of requests died somewhere — matching the observed 22% almost exactly. And the cause is named in the table: the search agent had started omitting source_urls from its brief after a change to its output schema. Nothing was down. One field went missing.
This is the general shape. End-to-end success is the product of the per-hop success rates, so a single hop at 78% ruins a chain of otherwise excellent links, and only per-pair measurement finds it.
Per-agent health checks that mean something
A health check returning {"status": "ok"} because the process is running is worse than no health check, because it produces confident green dashboards during an outage. Distinguish three levels:
| Level | Question | Used by | Cost |
|---|---|---|---|
| Liveness | Is the process alive and not deadlocked? | The orchestrator, to restart it | Microseconds |
| Readiness | Can it accept work right now? | The router, to decide where to send | Milliseconds |
| Capability | Can it actually do its job end to end? | Deploy gates and periodic probes | Seconds and tokens |
1class AgentHealth:2 def __init__(self, agent, max_queue=50):3 self.agent, self.max_queue = agent, max_queue4 self.last_success_ts = time.time()5 self.consecutive_failures = 067 def liveness(self) -> dict:8 # Only: is the loop turning? No external calls.9 stale = time.time() - self.agent.last_loop_tick10 return {"alive": stale < 30, "loop_stale_s": round(stale, 1)}1112 def readiness(self) -> dict:13 reasons = []14 if self.agent.queue_depth >= self.max_queue:15 reasons.append("queue full")16 if self.consecutive_failures >= 5:17 reasons.append("circuit open")18 if time.time() - self.last_success_ts > 600:19 reasons.append("no success in 10 minutes")20 return {"ready": not reasons, "reasons": reasons,21 "queue_depth": self.agent.queue_depth}2223 def capability(self) -> dict:24 # A real, tiny task through the real path. Run every few minutes.25 try:26 t0 = time.time()27 out = self.agent.run(CANARY_TASK, budget_tokens=400)28 ok = CANARY_EXPECTED in out.get("answer", "")29 return {"capable": ok, "latency_s": round(time.time() - t0, 2)}30 except Exception as exc:31 return {"capable": False, "error": f"{type(exc).__name__}: {exc}"}The capability check is the one that catches the interesting failures: an expired API key, a model deprecation, a tool whose upstream schema changed. All three leave liveness and readiness perfectly green. Keep the canary task small and deterministic — a fixed question with a known substring in the answer — and run it on a timer rather than per request, so it costs a few hundred tokens per agent per five minutes rather than doubling your bill.
Notice that readiness returns reasons. A boolean tells you an agent is not taking work; the reason tells you whether to scale it, restart it, or look upstream.
Tracing a request across agents, queues and retries
Metrics tell you something is wrong. Traces tell you where. The requirement is a single identifier that survives every hop, including hops through a queue and hops through a retry — which is exactly where naive tracing breaks, because a queue is a process boundary and most tracing libraries lose context across it.
1import uuid, contextvars, time23_ctx = contextvars.ContextVar("trace_ctx", default=None)45class Span:6 def __init__(self, name, kind, parent=None, trace_id=None):7 self.trace_id = trace_id or (parent.trace_id if parent8 else uuid.uuid4().hex)9 self.span_id = uuid.uuid4().hex[:16]10 self.parent_id = parent.span_id if parent else None11 self.name, self.kind = name, kind12 self.attrs: dict = {}13 self.start = time.time()14 self.end: float | None = None1516 def finish(self, **attrs):17 self.attrs.update(attrs)18 self.end = time.time()19 EXPORTER.export(self)2021 def carrier(self) -> dict:22 """Inject into a queue payload so the consumer can continue the trace."""23 return {"trace_id": self.trace_id, "parent_span_id": self.span_id}2425def enqueue_with_trace(queue, fn, payload, span: Span):26 payload = {**payload, "_trace": span.carrier()} # travels with the task27 return queue.enqueue(fn, payload)2829def run_task_with_trace(payload, agent_name):30 carrier = payload.pop("_trace", {})31 parent = _FakeParent(**carrier) if carrier else None32 span = Span(f"{agent_name}.run", kind="agent", parent=parent)33 try:34 result = AGENTS[agent_name].run(payload)35 span.finish(status="ok", tokens=result["tokens"],36 tool_calls=len(result["tool_calls"]))37 return result38 except Exception as exc:39 span.finish(status="error", error=type(exc).__name__)40 raiseThe _trace key inside the task payload is the whole trick. Trace context is normally propagated in HTTP headers; a queue has no headers, so you put it in the message body. Without this, every trace ends at the enqueue call and starts fresh in the worker, and you get dozens of one-span traces instead of one useful waterfall.
Here is what the resulting trace showed for a slow request in the incident:
trace 9f2c… total 41.8s├─ planner.run 1.9s tokens=2,140├─ handoff planner→search 0.1s├─ search.run 3.4s tokens=6,880 tool_calls=8├─ handoff search→verifier 22.6s ◄── queue wait, not agent time│ └─ attempt 1 declined reason="brief missing 'source_urls'"│ └─ attempt 2 declined reason="brief missing 'source_urls'"│ └─ attempt 3 accepted (fallback brief)├─ verifier.run 9.2s tokens=11,400└─ writer.run 4.6s tokens=8,020Total agent processing is 1.9 + 3.4 + 9.2 + 4.6 = 19.1 seconds. Total elapsed is 41.8. So 54% of the request's life was spent in one handoff, retrying against a validation failure. No per-agent dashboard could show that, because no agent was slow.
Attributes worth attaching to every agent span, because they are the ones you will wish you had: agent name and version, model name, prompt and completion tokens, number of tool calls, retry attempt number, delegation depth, and the decline reason where applicable. Tokens on the span are what let you answer "which agent is burning the budget" without a separate cost pipeline.
Alerting on coordination, not infrastructure
Infrastructure alerts — CPU, memory, process down — were all silent during the incident, and they were right to be. The alerts you need fire on relationships.
| Symptom | Likely cause | First fix |
|---|---|---|
| Handoff success rate for one pair drops below 90% | Brief schema changed; receiver at capacity | Read the top decline reason for that pair |
| p99 handoff latency > 10× p50 | Bimodal queueing behind one slow consumer | Check queue depth and worker count for the receiver |
| Re-delegation rate rises | Routing is choosing the wrong first agent | Inspect the router's classification distribution |
| Escalation rate rises with no error rate change | A guard is firing: depth cap, deadline, no consensus | Break escalations down by guard |
| Tokens per request rise, success rate flat | An agent is looping to its step cap | Look at tool_calls per span, p99 |
| Completion rate falls while all agents healthy | Results produced but never collected | Check result TTL and the collector's liveness |
Two rules make these alerts survivable. Alert on change, not thresholds: "handoff success for this pair fell more than 5 points versus the 7-day baseline" fires on real regressions and stays quiet during normal variation, whereas "below 95%" fires forever on a pair that legitimately sits at 94%. And require a minimum sample size — a pair with 3 handoffs in the window and 1 decline is at 67%, which is noise, and an alert that fires on it will be muted by the second week.
Where people get it wrong
Measuring agents instead of edges. The default instinct is per-service dashboards, because that is what the tooling gives you. Every failure in the opening incident lived on an edge. Add per-pair metrics or you cannot see them.
Averaging latency. A mean over a bimodal distribution describes nothing. It also hides exactly the failures users notice, because users notice the tail.
Shallow health checks. "The process responds to HTTP" is not health. An agent with a revoked API key answers that check instantly and cannot do a single unit of work.
Losing the trace at the queue. Traces that stop at enqueue turn one 42-second story into six unrelated fragments. Put the context in the payload.
Counting partial success as success. A report written from 3 of 4 research agents returns 200 and looks fine. Record the denominator: log how many contributors were expected and how many arrived, and alert when the ratio drops.
Sampling traces uniformly. At 1% sampling you will almost never capture the rare 30-second handoff, which is the one you need. Sample all errors, all slow requests, and 1% of the rest.
What this means when you build
Generate a trace ID at the system's edge — the HTTP handler, the webhook, the scheduler tick — and make it a required field on every envelope, task payload and state object. Not optional, not defaulted: required, so that code which forgets it fails at construction rather than producing an orphan trace you discover during an incident.
Instrument the edges before the nodes. Per-agent latency and token counts are easy and already partly covered by your model provider's dashboards. Per-pair handoff success, decline reasons and queue wait are the numbers nothing gives you for free, and they are the ones that explain a 22% failure rate that six green services cannot.
Build one view that answers "what happened to request X" in a single query. In practice that means one table where every row is one span with trace_id, parent_span_id, agent, phase, duration, status and tokens. It is unglamorous and it is the difference between diagnosing an incident in four minutes and correlating three log streams by eye for two hours.
And write the decline reason down every single time a handoff is refused. It costs one string. In the incident above, that one string — brief missing 'source_urls' — was the entire diagnosis, and without it the team had a 22% failure rate, six healthy services, and nowhere to start.