AutoGen Essentials

Course Content

AutoGen Essentials

7 sections · 28 lessons

How do you measure collaboration quality (handoffs, conflict resolution, redundancy)?


What you need to know

Why it matters

A team can reach the right answer while wasting half its turns: agents repeating each other, bouncing work back and forth, or one agent doing nothing useful. Outcome metrics do not show this; trace metrics do.

Metrics you can compute from the trace

MetricHowSignal
Mis-route rateSelected agent says "not my area", hands straight back, or its output is unusedBad descriptions or selector rules
RedundancyEmbedding similarity above about 0.9 to an earlier message; identical tool call and arguments twicePure waste; unclear roles
ContributionTokens per agent vs how often its content appears in the final answerAgents that cost but do not help
Conflict outcomeTag disagreements: resolved with evidence, by recency, or unresolved"By recency" means no real decider
Turn efficiencyTurns used vs the minimal path for that task typeDetours and loops
Critic valueScore with critic minus score withoutWhether reflection earns its cost

Computing redundancy and contribution

Python
from collections import Counterdef trace_metrics(result, final_text: str) -> dict:    msgs = [m for m in result.messages if m.source != "user"]    calls = [str(c.name) + c.arguments             for m in msgs if type(m).__name__ == "ToolCallRequestEvent"             for c in m.content]    dup_calls = sum(n - 1 for n in Counter(calls).values() if n > 1)    tokens = Counter()    for m in msgs:        if m.models_usage:            tokens[m.source] += m.models_usage.prompt_tokens + m.models_usage.completion_tokens    return {"duplicate_tool_calls": dup_calls, "tokens_by_agent": dict(tokens),            "turns": len(msgs)}

Add an embedding-based check for near-duplicate messages, and a small judge prompt to tag how conflicts ended.

The ablation

For each agent: remove it (or replace it with a pass-through), rerun the golden set, and compare success and cost. Keep agents whose removal lowers the score by more than noise; remove the rest.

A real-life example

A travel-planning team had five agents: planner, flights, hotels, budget and "local expert". Trace metrics over 300 runs showed:

  • local expert used 22% of tokens, but its suggestions appeared in only 9% of final plans.
  • 14% of runs had flights and planner both calling search_flights with the same arguments.
  • Budget disagreements ended "by recency" 60% of the time: whoever spoke last won.

Ablation on 120 golden tasks: without local expert, success went from 84% to 83% (within noise) and cost fell 21%, so they folded two of its prompts into planner. planner lost the search_flights tool. budget became the named decider on cost disputes. Duplicate tool calls dropped to 1%, and runs were 30% shorter.

Follow-up questions to expect

  • "Can an LLM judge collaboration quality?" — Yes, for tagging things like conflict outcomes, but check it against human labels on a sample and prefer code metrics where possible.
  • "What does a high mis-route rate usually mean?" — Overlapping roles or descriptions written for humans instead of for the selector.
  • "How often should you run ablations?" — After big changes to models, prompts or team shape; a stronger model often makes an agent unnecessary.