Course Content
Live Coding Interview Prep
7 sections · 50 lessons
Add reflection (self-correction) to an agent loop.
What you need to know
Reflection asks a model to review and improve its own output. It helps most with problems the model can recognise but did not avoid on the first pass: a missed requirement ("answer in 3 bullet points"), a format error, an internal contradiction.
It helps least with facts. A model that wrongly believes a fact will usually confirm it when asked to check. The fix is grounding the critique in something outside the model:
| signal | example | strength |
|---|---|---|
| Self-critique | "Check your answer for errors" | weak |
| Rubric critique | "Check: 3 bullets? cites a source? under 100 words?" | medium |
| Tool feedback | run the unit tests, validate the JSON schema, run the linter | strong |
| Retrieved evidence | compare claims with the source documents | strong |
Structure the critique as JSON so the loop can branch on it without guessing.
1import json2from collections.abc import Callable34CRITIC = (5 "You are a strict reviewer. Check the answer against the task for missed "6 "requirements, factual errors and unsupported claims. Reply with JSON only:\n"7 '{"passed": true, "issues": []}'8)910def reflect(task: str, llm_fn: Callable[[str], str], max_rounds: int = 2,11 check: Callable[[str], list[str]] | None = None) -> dict:12 """Generate, then critique and revise up to max_rounds times.1314 check(output) -> list of problems from an external tool (tests, schema); optional.15 """16 output = llm_fn(f"Task: {task}")17 reviews: list[dict] = []18 for _ in range(max_rounds):19 issues = check(output) if check else []20 if not issues:21 try:22 verdict = json.loads(llm_fn(f"{CRITIC}\n\nTask: {task}\n\nAnswer:\n{output}"))23 except json.JSONDecodeError:24 break # unreadable critique: keep what we have25 issues = [] if verdict.get("passed") else list(verdict.get("issues", []))26 reviews.append({"issues": issues})27 if not issues:28 break29 fixes = "\n".join(f"- {i}" for i in issues)30 output = llm_fn(f"Task: {task}\n\nPrevious answer:\n{output}\n\n"31 f"Rewrite it, fixing every issue below and changing nothing else:\n{fixes}")32 return {"output": output, "rounds": len(reviews), "reviews": reviews}The tricky parts:
- The external
checkruns first. If tests fail, their messages are the issues and no critic call is needed. The model's opinion is consulted only when the objective checks pass. - "Changing nothing else" in the revise prompt. Without it, the rewrite often fixes one issue and quietly breaks something that was right.
- An unparseable critique ends the loop with the current output instead of crashing.
passed: falsewith an emptyissueslist is treated as passed — there is nothing concrete to fix.
Complexity: worst case is 1 generation + max_rounds × (critique + revision) = 1 + 2 × max_rounds model calls, or fewer when check replaces the critic. Space is the kept reviews, O(rounds).
A real-life example
A task with a format rule, a scripted model and a real external check (count the bullet points):
1task = "List 3 benefits of UPI as bullet points."2replies = iter([3 "- Instant\n- Free", # first answer: only 2 bullets4 "- Instant transfers\n- No fees for users\n- Works 24x7", # revision5 '{"passed": true, "issues": []}', # critic, after checks pass6])7def bullets(out: str) -> list[str]:8 n = sum(line.startswith("- ") for line in out.splitlines())9 return [] if n == 3 else [f"expected 3 bullets, found {n}"]1011result = reflect(task, lambda _: next(replies), check=bullets)12print(result["rounds"], result["reviews"])13# 2 [{'issues': ['expected 3 bullets, found 2']}, {'issues': []}]14print(result["output"])15# - Instant transfers16# - No fees for users17# - Works 24x7| round | external check | critic called? | action |
|---|---|---|---|
| 1 | "expected 3 bullets, found 2" | no | revise with that issue |
| 2 | passes | yes → passed | stop |
Three model calls in total. The format problem was caught by 3 lines of Python, not by asking the model whether it had counted correctly.
Code-generation assistants use exactly this: generate, run the tests, feed failures back, and stop when the tests pass.
Follow-up questions to expect
- "Why not just use a stronger model instead?" — Often that is the better trade: one call to a stronger model can beat three calls to a weaker one. Reflection earns its place when there is an external check to ground it.
- "How do you know the critic is any good?" — Evaluate it on a labelled set of good and bad answers; measure how often it catches the bad ones and wrongly flags the good ones.
- "Can reflection make answers worse?" — Yes. Revisions can over-edit or add hedging. Keep the round count low and compare quality with and without reflection.