Course Content
Agentic AI Patterns
9 sections · 50 lessons
What is reflection in agentic workflows, and how does it improve reasoning?
What you need to know
Why it works, and when it does not
It works when the evaluator has information the generator lacked:
- A test runner output or compiler error.
- A schema validation failure.
- A calculation that does not match.
- A retrieved fact that contradicts the draft.
It is weak when the same model re-reads its own answer with nothing new. Research on self-correction has found that models often fail to fix their own reasoning errors without outside feedback, and can even change right answers to wrong ones. Every round also costs a full model call in latency and tokens.
Designing the loop
- Give the critic a rubric and require a structured verdict: pass or fail, a score, and a list of specific issues.
- Keep the critic's context separate from the generator's reasoning, so it is not anchored by it.
- Prefer an executable check to an LLM judge whenever one exists.
- Cap rounds (2 or 3) and stop when the score does not improve.
1def evaluator_optimizer(generate, evaluate, task, max_rounds=3):2 best, best_score, feedback = None, -1, None3 for round_no in range(1, max_rounds + 1):4 draft = generate(task, feedback)5 verdict = evaluate(task, draft) # {"pass": bool, "score": int, "issues": [...]}6 if verdict["pass"]:7 return draft, round_no, "passed"8 if verdict["score"] <= best_score:9 break # no improvement: stop paying10 best, best_score, feedback = draft, verdict["score"], verdict["issues"]11 return best, round_no, "needs_review"The loop returns the best draft seen, how many rounds it took, and a status. A needs_review status goes to a human queue, not to the customer.
A real-life example
A procurement agent writes a vendor-comparison summary for a buyer. The evaluator is mostly code, not a model:
- Does every vendor's total include GST? (regex plus arithmetic check)
- Does each total match the sum of line items? (computed)
- Is the delivery time stated for every vendor? (field check)
- Is the recommendation the lowest total that meets the delivery deadline? (computed)
Round 1: the draft quotes vendor B at Rs 1,18,000, but that excludes 18% GST. The evaluator returns issues: ["vendor B total excludes GST", "delivery time missing for vendor D"]. Round 2 fixes both and passes. About 70% of summaries pass in round 1, 25% in round 2, and 5% go to review.
A team member suggested adding a second LLM critic to "improve the writing style". In testing it changed correct numbers in 3 of 50 cases, so they dropped it.
Follow-up questions to expect
- "How is reflection different from self-consistency?" — Self-consistency samples several independent answers and votes. Reflection revises one answer using critique. Voting needs comparable final answers; reflection needs good criteria.
- "What is Reflexion?" — A 2023 method where the agent writes a short lesson after a failed attempt and keeps it in memory for the next attempt. It links reflection to episodic memory.
- "Do reasoning models still need reflection?" — Less for internal logic, since they already reconsider while thinking. External-signal reflection, like running tests, still helps a lot.