Course Content
Multi-Agent Systems and Collaboration
4 sections · 12 lessons
Delegation and Task Handoff
An insurance company built a claims assistant. The triage agent read the claim, pulled the policy, fetched the photos, checked the repair estimate, and talked to the customer. When something looked suspicious it handed the claim to a fraud agent. The handoff was one line:
fraud_agent.run({"claim_id": 4471})The fraud agent then re-fetched the policy, re-fetched all eleven photos, re-read the estimate, and asked the customer three questions they had already answered eight minutes earlier. Average handoff added 22 seconds of latency and 31,000 tokens. In a sample of 400 handoffs, 38% contained at least one question the customer had already answered.
The engineers' fix was to pass the whole conversation instead. That solved the re-asking and created a new problem: the fraud agent's context now contained the triage agent's chatty reassurances and its speculative reasoning ("this looks routine to me"), and the fraud agent started agreeing with conclusions it was supposed to independently check. Fraud detection rate fell by a fifth.
Neither the thin handoff nor the fat one was right. What was missing was a deliberate decision about what crosses the boundary.
Delegation is not the supervisor pattern
These get conflated constantly, and the difference determines how you write the code.
| Supervisor-worker | Delegation | |
|---|---|---|
| Who initiates | The supervisor, always | Any agent, when it hits its own limit |
| Relationship | Fixed hierarchy, defined up front | Peer to peer, decided at runtime |
| The delegator's role after | Tracks every task to completion | May hand over entirely and stop caring |
| Can the receiver refuse? | No — it is an assignment | Yes — it is an offer |
| Nesting | Rare; workers do not sub-assign | Common; delegates re-delegate |
| Typical trigger | "There are 2,000 items to process" | "This one item is not mine to handle" |
The triage agent was not a supervisor. It did not have a work queue and it was not managing the fraud agent. It reached a point where the task exceeded its authority and competence, and it passed the task on. That is delegation, and the useful mental model is a colleague saying "this is really one for the fraud team" — including the possibility that the fraud team says "not ours, actually".
An assignment is given. A delegation is offered, and an offer that cannot be declined is just an assignment with worse error handling.
The delegated task and its life cycle
A delegated task passes through more states than an assigned one, because the receiver has a say and because delegation can nest.
PROPOSED ──accept──► ACCEPTED ──► IN_PROGRESS ──► RETURNED │ │ ├──decline──► REJECTED ├──► ESCALATED (needs a human) │ │ └──timeout──► EXPIRED └──► SUB_DELEGATED1from dataclasses import dataclass, field2from enum import Enum3import time, uuid45class DState(str, Enum):6 PROPOSED = "proposed"; ACCEPTED = "accepted"; REJECTED = "rejected"7 EXPIRED = "expired"; IN_PROGRESS = "in_progress"8 RETURNED = "returned"; ESCALATED = "escalated"910@dataclass11class Delegation:12 goal: str # what success looks like, in words13 from_agent: str14 to_agent: str15 brief: dict # the curated context (see below)16 deliverable: dict # required shape of the answer17 id: str = field(default_factory=lambda: str(uuid.uuid4()))18 root_id: str | None = None # the original request19 chain: list[str] = field(default_factory=list) # agents already involved20 depth: int = 021 token_budget: int = 50_00022 deadline_ts: float = field(default_factory=lambda: time.time() + 120)23 state: DState = DState.PROPOSED24 result: dict | None = None25 decline_reason: str | None = NoneFour of those fields are pure defence and they are the ones beginners omit: chain prevents cycles, depth prevents infinite nesting, token_budget prevents a delegate from spending the whole request's allowance, and deadline_ts prevents a silent hang. Each corresponds to a failure that is very hard to diagnose after the fact and trivial to prevent here.
The delegating agent
Delegation has four steps, and skipping any one of them produces a recognisable bug.
1class DelegatingAgent:2 def __init__(self, name, registry, max_depth=3):3 self.name, self.registry, self.max_depth = name, registry, max_depth45 # 1. Decide whether to delegate at all6 def should_delegate(self, task) -> str | None:7 needed = self.capability_required(task) # e.g. "fraud_review"8 if needed in self.own_capabilities:9 return None # do it yourself10 return needed1112 # 2. Choose a delegate13 def choose(self, capability, exclude: set[str]) -> str | None:14 candidates = [a for a in self.registry.providing(capability)15 if a.name not in exclude and a.accepting()]16 if not candidates:17 return None18 return min(candidates, key=lambda a: a.queue_depth).name1920 # 3. Build the brief, 4. Offer it21 def delegate(self, task, parent: Delegation | None = None):22 cap = self.should_delegate(task)23 if cap is None:24 return self.execute(task)2526 depth = 0 if parent is None else parent.depth + 127 chain = [] if parent is None else list(parent.chain)28 if depth > self.max_depth:29 raise DelegationTooDeep(chain)3031 target = self.choose(cap, exclude=set(chain) | {self.name})32 if target is None:33 return self.escalate(task, reason=f"no available {cap} agent")3435 d = Delegation(36 goal=task.goal, from_agent=self.name, to_agent=target,37 brief=self.build_brief(task), deliverable=task.deliverable,38 root_id=(parent.root_id if parent else None),39 chain=chain + [self.name], depth=depth,40 token_budget=(parent.token_budget // 2 if parent else 50_000),41 )42 return self.registry.get(target).receive(d)Note exclude=set(chain) | {self.name} in step 2. That single expression is what stops a claim being handed from triage to fraud to triage forever: an agent already in the chain cannot be chosen again.
Confirming the handoff
Delegation must be two-phase — offer, then accept or decline — because the receiver has information the sender lacks. It may be at capacity, it may be missing a tool, it may be able to see that the brief is incomplete.
1class ReceivingAgent:2 def receive(self, d: Delegation):3 if time.time() > d.deadline_ts:4 d.state = DState.EXPIRED5 return d6 ok, reason = self.can_accept(d)7 if not ok:8 d.state, d.decline_reason = DState.REJECTED, reason9 return d # sender must handle this10 d.state = DState.ACCEPTED11 return self.run(d)1213 def can_accept(self, d: Delegation) -> tuple[bool, str | None]:14 if self.queue_depth >= self.max_queue:15 return False, "at capacity"16 missing = [k for k in self.required_brief_keys if k not in d.brief]17 if missing:18 return False, f"brief missing: {missing}"19 if d.token_budget < self.min_viable_budget:20 return False, "budget too small for a reliable answer"21 return True, NoneThe brief missing check is worth its weight. Without it, an incomplete brief becomes a bad answer — the fraud agent guesses, produces something plausible, and nobody knows the guess happened. With it, the incompleteness surfaces immediately as a decline the sender can act on, which turns a quality bug into a control-flow event.
And the sender must actually handle a decline. Three reasonable responses: try the next-best delegate, enrich the brief and re-offer, or escalate. What you must not do is treat a decline as a failure of the whole request — a decline is normal.
Context preservation: the heart of it
There are exactly three things you can send across a handoff, and the choice has a large, measurable cost.
| Approach | What crosses | Tokens per handoff | Risk |
|---|---|---|---|
| Full transcript | Every message so far | ~18,400 | Anchoring on the sender's conclusions; cost grows per hop |
| Structured brief | Curated facts, constraints, deliverable | ~940 | Omitting something that mattered |
| Reference only | An ID into shared storage | ~40 | Receiver must know what to fetch; latency of fetching |
Run the arithmetic over a three-hop chain. Full transcript is not merely 3 × 18,400 — it grows, because each agent appends its own reasoning: roughly 18,400 + 23,000 + 28,500 = 69,900 tokens of pure handoff. Structured briefs cost about 940 each, so 2,820 in total: a 25-fold reduction. That is the difference between a handoff being free and being the dominant cost of your system.
But the transcript approach also caused the fraud-detection drop, and that is the more interesting failure. When the fraud agent read "this looks routine to me" in its context, it was being primed. An independent check that inherits the first agent's conclusion is not independent. Sometimes you should deliberately withhold the sender's reasoning.
A handoff brief is an editorial decision, not a serialisation problem. What you leave out is as deliberate as what you put in.
What a good brief contains
1def build_brief(self, task) -> dict:2 return {3 # 1. The user's actual words, never a paraphrase4 "original_request": task.user_text,56 # 2. Facts established, each with where it came from7 "established": {8 "policy_number": {"value": "PX-88213",9 "source": "policy_db", "confidence": 1.0},10 "incident_date": {"value": "2026-03-14",11 "source": "customer", "confidence": 0.9},12 "estimate_gbp": {"value": 4310,13 "source": "garage_pdf", "confidence": 1.0},14 },1516 # 3. Hard constraints the delegate must respect17 "constraints": ["customer is travelling until the 22nd",18 "policy excludes flood damage"],1920 # 4. Pointers to bulk data rather than the data itself21 "artifacts": {"photos": "s3://claims/4471/photos/",22 "estimate_pdf": "s3://claims/4471/est.pdf"},2324 # 5. What has already been tried, so it is not repeated25 "already_done": ["policy fetched", "customer contacted 08:42",26 "photos OCR'd"],2728 # 6. Explicitly NOT the sender's opinion of the answer29 }Each section earns its place. original_request verbatim, because paraphrases lose the nuance that turns out to matter. source and confidence on every fact, because the delegate needs to know which claims are verified and which are the customer's assertion. already_done, because it directly removes the 22 seconds of re-fetching. artifacts as pointers, because eleven photos do not belong in a prompt.
The hybrid this produces — a structured brief with references for bulk data — is the shape that works in practice. It costs about 940 tokens, preserves everything that matters, and lets the delegate fetch the two photos it actually needs rather than inheriting all eleven.
Sub-delegation
The fraud agent looks at claim 4471, decides the estimate itself is the suspicious part, and wants a specialist in repair pricing. It delegates onward. This is correct behaviour and it introduces three new hazards.
Depth. Each hop adds latency and loses context. A cap of three is right for most systems; beyond that, the brief has been filtered so many times that the last agent is working from a summary of a summary. Enforce it in code, as DelegationTooDeep above, not by hoping.
Budget. A delegate that inherits the full token budget can spend all of it, leaving nothing for the parent to finish the job. Halving at each hop is a simple, effective rule: 50,000 at the root, 25,000 at depth 1, 12,500 at depth 2. Combined with the min_viable_budget check in can_accept, a chain that has become too deep to fund declines rather than producing a starved, low-quality answer.
Provenance. When the answer comes back three hops later, the root caller needs to know who touched it. The chain field gives you that for free, and it turns "the fraud verdict is wrong" into "the pricing specialist at depth 2 used the wrong regional rate table".
1def unwind(d: Delegation) -> dict:2 """Return the result annotated with its full delegation path."""3 return {"result": d.result,4 "path": " -> ".join(d.chain + [d.to_agent]),5 "depth": d.depth,6 "tokens_allocated": d.token_budget}7# {"result": {...}, "path": "triage -> fraud -> pricing", "depth": 2, ...}Preventing circular and unbounded delegation
The classic bug: triage decides the claim is fraud's, fraud decides the ambiguity is triage's, and the two hand it back and forth. Each decision is individually defensible. Together they never terminate, and because each hop looks like normal activity, monitoring shows a healthy busy system.
Four independent guards, and you want all four because each catches a different shape of the problem:
| Guard | Catches | Implementation |
|---|---|---|
| Chain exclusion | A returns to an agent already in the path | exclude=set(chain) when choosing |
| Depth cap | Ever-deeper nesting that never repeats an agent | depth > max_depth raises |
| Budget decay | Chains that are legal but wasteful | Halve the budget each hop |
| Root deadline | Everything else, including slow-but-legal chains | One wall-clock deadline carried from the root |
Chain exclusion alone is insufficient because A can delegate to B, B to C, C to D — never repeating, never terminating. The depth cap alone is insufficient because a two-hop chain can still take four minutes. The root deadline is the backstop that catches everything, and it must be an absolute timestamp set once at the root, not a per-hop timeout: three hops with a 120-second per-hop timeout allows six minutes.
When a guard fires, the right response is almost never to fail the request. It is to escalate — return to the caller with the partial results gathered so far, the chain that was tried, and a clear reason. A human reading "triage → fraud → pricing, depth limit reached, no verdict, here is what each found" can act. A human reading DelegationTooDeep cannot.
Where people get it wrong
"Passing the whole conversation is the safe default." It is the expensive default and sometimes the wrong one. It costs 20 times as much, grows at every hop, and it anchors the receiver on the sender's conclusions — which is fatal precisely when the receiver's job is independent verification.
"If the delegate is capable, it will accept." Capability is one of at least three reasons to decline. Capacity and budget are the others, and an agent that cannot decline for capacity reasons will simply queue work invisibly until its latency is measured in minutes.
"Delegation is fire-and-forget." Sometimes it genuinely is — a full handoff where the delegate owns the outcome and replies to the user directly. But if the delegator needs the result to continue, it must track the delegation with a deadline, or a delegate that dies silently leaves the original request hanging forever with no error anywhere.
"The brief is whatever the sender's state happens to be." This is how the {"claim_id": 4471} handoff happens: nobody designed the brief, so it became whatever variable was nearest. Define the brief as a schema, and have the receiver validate it, and the problem cannot occur.
"Confidence scores are noise, just send the values." A fact asserted by the customer and a fact read from the policy database are not the same kind of fact. When the delegate cannot distinguish them, it will treat an unverified claim as ground truth — and the whole point of handing to a fraud agent is that unverified claims are the subject matter.
What this means when you build
Write the brief schema before you write either agent. It is the interface, and like any interface it should be specified rather than emergent. Concretely: define the dataclass, list which keys are required by each receiving capability, and write a test that constructs a brief from a realistic sender state and asserts that a receiver accepts it. That test catches the {"claim_id": 4471} class of bug before it reaches a customer.
Log every delegation as a record, not a log line: id, root_id, chain, depth, state transitions with timestamps, tokens spent, and the decline reason where applicable. With that table you can answer the questions that actually come up — which handoffs are declined most often and why, what the median chain depth is, which agent pair produces the most re-asked questions. Without it you are debugging a distributed control flow from three unconnected log streams.
Set the four guards on day one, even for a two-agent system. They cost about fifteen lines. Retrofitting them after a production incident means retrofitting them into every delegation site you have written since, and the incident you are retrofitting after is one where an agent pair spent eleven hours and a large amount of money handing one claim back and forth.
Finally, make the decline path a first-class citizen in your tests. Most teams test the happy path — A delegates, B accepts, B returns a result — and discover on the first busy day that nothing handles a decline, so a declined delegation returns None, which the caller treats as an answer.