Course Content
AutoGen Essentials
7 sections · 28 lessons
How do you evaluate a multi-agent AutoGen system (success rate, tool accuracy, latency, cost)?
What you need to know
The metric set
| Metric | How to measure | Why it matters |
|---|---|---|
| Task success rate | Golden tasks, scored by code (tests pass, correct value) or a rubric judge | The number users feel |
| Tool accuracy | Per call: right tool? valid args? succeeded? Over- and under-calls separately | Wrong calls cost money or cause side effects |
| Cost per task | Tokens and money per run; report p50 and p95 | Multi-agent cost has a long tail where loops hide |
| Latency | End-to-end p50 and p95; turns per task | Turns, not single calls, dominate multi-agent latency |
| Termination health | Share of runs ending on the semantic condition vs MaxMessageTermination or timeout | A rising cap-hit rate means loops |
| Stability | Same task run 3 to 5 times: how often does the outcome change? | Tells you if a pass was luck |
| Human intervention | Share of runs escalated or corrected | Real operating cost |
Where the numbers come from in AutoGen 0.4+
Python
1result = await team.run(task=case["input"])23agents_spoke = [m.source for m in result.messages]4tool_calls = [m for m in result.messages if type(m).__name__ == "ToolCallRequestEvent"]5tokens = sum((m.models_usage.prompt_tokens + m.models_usage.completion_tokens)6 for m in result.messages if m.models_usage)78record = {9 "id": case["id"],10 "stop_reason": result.stop_reason, # which condition ended the run11 "turns": len(agents_spoke),12 "tool_calls": len(tool_calls),13 "tokens": tokens,14 "passed": case["check"](result), # your scorer15}- Messages carry
models_usagewhere a model call produced them, so you can attribute tokens per agent. The model client'stotal_usage()gives totals for that client. Calls made by the team itself, such as theSelectorGroupChatspeaker selection, may not show up on agent messages, so read those from that model client's usage or from traces. - AutoGen's runtime emits OpenTelemetry spans; with a tracer provider configured you get per-agent and per-tool spans in Jaeger, Azure Monitor or similar.
Evaluating the right way round
- Fix the golden set first (see the next lesson).
- Run it on every change to prompts, tools, models or team shape.
- Compare to the last release on success and cost. A 2-point quality gain that costs 4 times as much per task is usually not shippable.
A real-life example
A bank's customer-support triage team (triage, cards, loans, human handoff) was about to switch to a newer, cheaper model. The team ran 150 golden tasks, three times each:
| Old model | New model | |
|---|---|---|
| Success rate | 87% | 89% |
| Wrong-tool calls | 4.1% | 2.9% |
| Median cost per task | ₹2.40 | ₹1.10 |
| p95 cost per task | ₹7.90 | ₹12.60 |
| Cap-hit runs | 1.3% | 4.0% |
The averages looked great, but the p95 cost and cap-hit rate showed a new loop: the new model often handed off between cards and loans on credit-card EMI questions. They clarified both descriptions and added a hop cap; the cap-hit rate fell to 1.1% and p95 cost to ₹5.20 before rollout.
Follow-up questions to expect
- "How do you score open-ended answers?" — With an LLM judge using a written rubric and reference facts, validated against human labels on a sample, and a different model from the one being tested where possible.
- "Why p95 and not just the average?" — Multi-agent runs have a long tail: a few looping runs cost 10 times the median. Averages hide them.
- "Offline or online evaluation?" — Both: golden sets before release, and in production track success signals (resolution, escalation, thumbs down), cost and cap hits per day.