Course Content
AutoGen Essentials
7 sections · 28 lessons
How do you build “golden tasks” and regression tests for agent workflows?
What you need to know
What a golden task looks like
1id: refund-outside-window2input: "I want a refund for order 8812, bought 90 days ago."3fixtures:4 orders: fixtures/orders_8812.json # delivered 90 days ago, ₹2,4995expect:6 tools_called: [lookup_order]7 tools_not_called: [issue_refund]8 final_mentions: ["30-day"]9 stop_reason_contains: "APPROVED"10 max_turns: 811 max_tokens: 20000Rules that keep them useful
- From reality. Sample real conversations, keep the interesting ones, remove personal data. Every production bug becomes a case.
- Invariants, not wording. Exact-string checks on LLM prose fail randomly. Check tools, side effects, schema, key facts.
- Test the negatives. "Did not refund", "did not reveal another customer's data" catch the most expensive failures.
- Stub the world. Tools read fixtures, so you test your agents, not your vendors' uptime. Keep a small live suite that runs nightly.
- Budgets are tests. Fail the build if turns or tokens per task jump.
- Handle non-determinism. Pin the model version; run flaky cases 3 times and require, for example, 3 of 3 for safety cases and 2 of 3 for quality cases.
Two levels of test
| Level | What it tests | Model calls |
|---|---|---|
| Wiring tests | Team routing, termination, tool plumbing | None: use ReplayChatCompletionClient from autogen_ext.models.replay, which returns scripted replies |
| Golden tasks | Real behaviour of prompts plus model | Real model, stubbed tools |
1from autogen_core.models import ModelInfo2from autogen_ext.models.replay import ReplayChatCompletionClient34fake = ReplayChatCompletionClient(5 ["Checking the order.", "Outside the 30-day window. APPROVED"],6 model_info=ModelInfo(vision=False, function_calling=True, json_output=False,7 family="unknown", structured_output=False),8)9team = build_support_team(model_client=fake, tools=stub_tools("orders_8812.json"))10result = await team.run(task="Refund order 8812")11assert "APPROVED" in result.stop_reasonThe replay client returns its scripted replies in order, and function_calling=True is needed because the default replay client refuses agents with tools. Wiring tests run in seconds and cost nothing, so they catch broken termination or routing before you spend tokens on golden tasks.
A real-life example
A food-delivery company's customer-support triage team had a bad release: a prompt tweak to make the refunds agent "more empathetic" also made it approve refunds for orders delivered on time. It cost about ₹3 lakh in two days before anyone noticed.
They built a suite of 64 golden tasks from real chats, including 12 "must not refund" cases. Each case stubs lookup_order from a fixture and records every tool call. In CI, every PR touching prompts, tools or the model version runs the suite; safety cases must pass 3 of 3 runs. Six weeks later, a model upgrade failed two "must not refund" cases in CI, and the team fixed the prompt before release instead of after.
Follow-up questions to expect
- "How many golden tasks do you need?" — Start with 30 to 50 covering main intents, edge cases and past bugs; grow it as bugs appear. Coverage of failure types matters more than size.
- "How do you keep them from going stale?" — Review quarterly, refresh fixtures when data shapes change, and retire cases for features that no longer exist.
- "Temperature zero makes it deterministic, right?" — Not fully; providers can still vary, and some reasoning models do not accept a temperature setting. Pin the model version and use repeated runs with pass rates.