Course Content
AutoGen Essentials
7 sections · 28 lessons
How do you measure collaboration quality (handoffs, conflict resolution, redundancy)?
What you need to know
Why it matters
A team can reach the right answer while wasting half its turns: agents repeating each other, bouncing work back and forth, or one agent doing nothing useful. Outcome metrics do not show this; trace metrics do.
Metrics you can compute from the trace
| Metric | How | Signal |
|---|---|---|
| Mis-route rate | Selected agent says "not my area", hands straight back, or its output is unused | Bad descriptions or selector rules |
| Redundancy | Embedding similarity above about 0.9 to an earlier message; identical tool call and arguments twice | Pure waste; unclear roles |
| Contribution | Tokens per agent vs how often its content appears in the final answer | Agents that cost but do not help |
| Conflict outcome | Tag disagreements: resolved with evidence, by recency, or unresolved | "By recency" means no real decider |
| Turn efficiency | Turns used vs the minimal path for that task type | Detours and loops |
| Critic value | Score with critic minus score without | Whether reflection earns its cost |
Computing redundancy and contribution
1from collections import Counter23def trace_metrics(result, final_text: str) -> dict:4 msgs = [m for m in result.messages if m.source != "user"]5 calls = [str(c.name) + c.arguments6 for m in msgs if type(m).__name__ == "ToolCallRequestEvent"7 for c in m.content]8 dup_calls = sum(n - 1 for n in Counter(calls).values() if n > 1)9 tokens = Counter()10 for m in msgs:11 if m.models_usage:12 tokens[m.source] += m.models_usage.prompt_tokens + m.models_usage.completion_tokens13 return {"duplicate_tool_calls": dup_calls, "tokens_by_agent": dict(tokens),14 "turns": len(msgs)}Add an embedding-based check for near-duplicate messages, and a small judge prompt to tag how conflicts ended.
The ablation
For each agent: remove it (or replace it with a pass-through), rerun the golden set, and compare success and cost. Keep agents whose removal lowers the score by more than noise; remove the rest.
A real-life example
A travel-planning team had five agents: planner, flights, hotels, budget and "local expert". Trace metrics over 300 runs showed:
local expertused 22% of tokens, but its suggestions appeared in only 9% of final plans.- 14% of runs had
flightsandplannerboth callingsearch_flightswith the same arguments. - Budget disagreements ended "by recency" 60% of the time: whoever spoke last won.
Ablation on 120 golden tasks: without local expert, success went from 84% to 83% (within noise) and cost fell 21%, so they folded two of its prompts into planner. planner lost the search_flights tool. budget became the named decider on cost disputes. Duplicate tool calls dropped to 1%, and runs were 30% shorter.
Follow-up questions to expect
- "Can an LLM judge collaboration quality?" — Yes, for tagging things like conflict outcomes, but check it against human labels on a sample and prefer code metrics where possible.
- "What does a high mis-route rate usually mean?" — Overlapping roles or
descriptions written for humans instead of for the selector. - "How often should you run ablations?" — After big changes to models, prompts or team shape; a stronger model often makes an agent unnecessary.