Agentic AI Patterns

Course Content

Agentic AI Patterns

9 sections · 50 lessons

What is reflection in agentic workflows, and how does it improve reasoning?


Evaluator-optimizer on a vendor comparisondraft the summarycode checksGST, sums, datespass? return itrevise with thelisted issuesnewinformationmax 3 roundsRound 1 missed 18 percent GST on vendor B; round 2 passed.
Reflection converges when the critic knows something the generator did not — a re-read with nothing new mostly repeats itself.

What you need to know

Why it works, and when it does not

It works when the evaluator has information the generator lacked:

  • A test runner output or compiler error.
  • A schema validation failure.
  • A calculation that does not match.
  • A retrieved fact that contradicts the draft.

It is weak when the same model re-reads its own answer with nothing new. Research on self-correction has found that models often fail to fix their own reasoning errors without outside feedback, and can even change right answers to wrong ones. Every round also costs a full model call in latency and tokens.

Designing the loop

  • Give the critic a rubric and require a structured verdict: pass or fail, a score, and a list of specific issues.
  • Keep the critic's context separate from the generator's reasoning, so it is not anchored by it.
  • Prefer an executable check to an LLM judge whenever one exists.
  • Cap rounds (2 or 3) and stop when the score does not improve.
Python
def evaluator_optimizer(generate, evaluate, task, max_rounds=3):    best, best_score, feedback = None, -1, None    for round_no in range(1, max_rounds + 1):        draft = generate(task, feedback)        verdict = evaluate(task, draft)   # {"pass": bool, "score": int, "issues": [...]}        if verdict["pass"]:            return draft, round_no, "passed"        if verdict["score"] <= best_score:            break                          # no improvement: stop paying        best, best_score, feedback = draft, verdict["score"], verdict["issues"]    return best, round_no, "needs_review"

The loop returns the best draft seen, how many rounds it took, and a status. A needs_review status goes to a human queue, not to the customer.

A real-life example

A procurement agent writes a vendor-comparison summary for a buyer. The evaluator is mostly code, not a model:

  • Does every vendor's total include GST? (regex plus arithmetic check)
  • Does each total match the sum of line items? (computed)
  • Is the delivery time stated for every vendor? (field check)
  • Is the recommendation the lowest total that meets the delivery deadline? (computed)

Round 1: the draft quotes vendor B at Rs 1,18,000, but that excludes 18% GST. The evaluator returns issues: ["vendor B total excludes GST", "delivery time missing for vendor D"]. Round 2 fixes both and passes. About 70% of summaries pass in round 1, 25% in round 2, and 5% go to review.

A team member suggested adding a second LLM critic to "improve the writing style". In testing it changed correct numbers in 3 of 50 cases, so they dropped it.

Follow-up questions to expect

  • "How is reflection different from self-consistency?" — Self-consistency samples several independent answers and votes. Reflection revises one answer using critique. Voting needs comparable final answers; reflection needs good criteria.
  • "What is Reflexion?" — A 2023 method where the agent writes a short lesson after a failed attempt and keeps it in memory for the next attempt. It links reflection to episodic memory.
  • "Do reasoning models still need reflection?" — Less for internal logic, since they already reconsider while thinking. External-signal reflection, like running tests, still helps a lot.