AutoGen Essentials

Course Content

AutoGen Essentials

7 sections · 28 lessons

How do you evaluate a multi-agent AutoGen system (success rate, tool accuracy, latency, cost)?


The cheaper model, before a hop cap87%89%Rs 2.40Rs 1.10Rs 7.90Rs 12.601.3%4.0%Old modelNew modelSuccessMedian costp95 costCap-hit runsA cards-to-loans handoff loop hid in the tail.
The averages said ship it; the p95 cost and cap-hit rate showed a new loop that only a tail metric could see.

What you need to know

The metric set

MetricHow to measureWhy it matters
Task success rateGolden tasks, scored by code (tests pass, correct value) or a rubric judgeThe number users feel
Tool accuracyPer call: right tool? valid args? succeeded? Over- and under-calls separatelyWrong calls cost money or cause side effects
Cost per taskTokens and money per run; report p50 and p95Multi-agent cost has a long tail where loops hide
LatencyEnd-to-end p50 and p95; turns per taskTurns, not single calls, dominate multi-agent latency
Termination healthShare of runs ending on the semantic condition vs MaxMessageTermination or timeoutA rising cap-hit rate means loops
StabilitySame task run 3 to 5 times: how often does the outcome change?Tells you if a pass was luck
Human interventionShare of runs escalated or correctedReal operating cost

Where the numbers come from in AutoGen 0.4+

Python
result = await team.run(task=case["input"])agents_spoke = [m.source for m in result.messages]tool_calls = [m for m in result.messages if type(m).__name__ == "ToolCallRequestEvent"]tokens = sum((m.models_usage.prompt_tokens + m.models_usage.completion_tokens)             for m in result.messages if m.models_usage)record = {    "id": case["id"],    "stop_reason": result.stop_reason,     # which condition ended the run    "turns": len(agents_spoke),    "tool_calls": len(tool_calls),    "tokens": tokens,    "passed": case["check"](result),       # your scorer}
  • Messages carry models_usage where a model call produced them, so you can attribute tokens per agent. The model client's total_usage() gives totals for that client. Calls made by the team itself, such as the SelectorGroupChat speaker selection, may not show up on agent messages, so read those from that model client's usage or from traces.
  • AutoGen's runtime emits OpenTelemetry spans; with a tracer provider configured you get per-agent and per-tool spans in Jaeger, Azure Monitor or similar.

Evaluating the right way round

  1. Fix the golden set first (see the next lesson).
  2. Run it on every change to prompts, tools, models or team shape.
  3. Compare to the last release on success and cost. A 2-point quality gain that costs 4 times as much per task is usually not shippable.

A real-life example

A bank's customer-support triage team (triage, cards, loans, human handoff) was about to switch to a newer, cheaper model. The team ran 150 golden tasks, three times each:

Old modelNew model
Success rate87%89%
Wrong-tool calls4.1%2.9%
Median cost per task₹2.40₹1.10
p95 cost per task₹7.90₹12.60
Cap-hit runs1.3%4.0%

The averages looked great, but the p95 cost and cap-hit rate showed a new loop: the new model often handed off between cards and loans on credit-card EMI questions. They clarified both descriptions and added a hop cap; the cap-hit rate fell to 1.1% and p95 cost to ₹5.20 before rollout.

Follow-up questions to expect

  • "How do you score open-ended answers?" — With an LLM judge using a written rubric and reference facts, validated against human labels on a sample, and a different model from the one being tested where possible.
  • "Why p95 and not just the average?" — Multi-agent runs have a long tail: a few looping runs cost 10 times the median. Averages hide them.
  • "Offline or online evaluation?" — Both: golden sets before release, and in production track success signals (resolution, escalation, thumbs down), cost and cap hits per day.