Course Content
AutoGen Essentials
7 sections · 28 lessons
How can you add reflection or self-critique loops without wasting tokens or overfitting feedback?
What you need to know
Cost math
If the draft costs 4,000 tokens and each critique plus revision costs 5,000, then two rounds turn a 4,000-token task into 14,000 tokens. That is worth it for a legal summary, not for a weather answer.
Keep it cheap
- Gate it. Reflect only on high-value tasks, or when a cheap check fails (tests red, JSON invalid, low retrieval score).
- Verifier before critic. Running tests, validating a schema or re-running SQL is almost free and never hallucinates. Use an LLM critic only for what code cannot check.
- Bounded critique. "Reply
APPROVED, or list at most 3 issues that change correctness. No style comments." - Hard cap. Two revisions, then ship the best version with notes. Enforce with a termination condition.
- Cheaper critic model, and send a diff rather than the whole document on later rounds.
1critic = AssistantAgent("critic", model_client=small_client,2 system_message="Compare the latest code with the ORIGINAL task and the test "3 "output. Reply APPROVED, or at most 3 numbered bugs. No style.")4team = RoundRobinGroupChat([author, tests, critic],5 termination_condition=TextMentionTermination("APPROVED", sources=["critic"])6 | MaxMessageTermination(9)) # three rounds of 3 agentsHere tests is a CodeExecutorAgent that runs the test suite before the critic speaks, so the critic sees real results.
Overfitting to feedback
The writer starts pleasing the critic instead of the user: adding caveats the critic likes, drifting from the request. Guards:
- The critic scores against the original requirements, not its own earlier comments.
- A head-and-tail context keeps the original task in every revision.
- Track the final score against the user's criteria, not "did the critic approve".
A real-life example
A code-review agent pair generates fixes for failing unit tests in a large Django codebase. Without reflection, 58% of fixes passed the tests. With an unbounded LLM critic, 66% passed, but cost tripled and 12% of runs looped on naming nitpicks.
The redesign ran the tests first; only if they failed did the critic see the failure and the diff, with the three-issue contract, on a smaller model, capped at two rounds. The pass rate rose to 71%, because the critic now reasoned about real test output, and average cost was 1.6 times the no-reflection baseline instead of 3 times. The team checked monthly: round two still added about 4 points, so they kept it.
Follow-up questions to expect
- "Self-reflection or a separate critic?" — A separate agent with a different prompt, and ideally different evidence such as test output, catches more than the same model re-reading its own work.
- "How do you know reflection helps?" — Run the eval set with and without it and compare quality and cost per task.
- "What if the critic is wrong?" — Give the writer permission to reject a point with a reason, and let the verifier (tests) or a named decider settle it.