Live Coding Interview Prep

Course Content

Live Coding Interview Prep

7 sections · 50 lessons

Add reflection (self-correction) to an agent loop.


What you need to know

Reflection asks a model to review and improve its own output. It helps most with problems the model can recognise but did not avoid on the first pass: a missed requirement ("answer in 3 bullet points"), a format error, an internal contradiction.

It helps least with facts. A model that wrongly believes a fact will usually confirm it when asked to check. The fix is grounding the critique in something outside the model:

signalexamplestrength
Self-critique"Check your answer for errors"weak
Rubric critique"Check: 3 bullets? cites a source? under 100 words?"medium
Tool feedbackrun the unit tests, validate the JSON schema, run the linterstrong
Retrieved evidencecompare claims with the source documentsstrong

Structure the critique as JSON so the loop can branch on it without guessing.

Python
import jsonfrom collections.abc import CallableCRITIC = (    "You are a strict reviewer. Check the answer against the task for missed "    "requirements, factual errors and unsupported claims. Reply with JSON only:\n"    '{"passed": true, "issues": []}')def reflect(task: str, llm_fn: Callable[[str], str], max_rounds: int = 2,            check: Callable[[str], list[str]] | None = None) -> dict:    """Generate, then critique and revise up to max_rounds times.    check(output) -> list of problems from an external tool (tests, schema); optional.    """    output = llm_fn(f"Task: {task}")    reviews: list[dict] = []    for _ in range(max_rounds):        issues = check(output) if check else []        if not issues:            try:                verdict = json.loads(llm_fn(f"{CRITIC}\n\nTask: {task}\n\nAnswer:\n{output}"))            except json.JSONDecodeError:                break                                  # unreadable critique: keep what we have            issues = [] if verdict.get("passed") else list(verdict.get("issues", []))        reviews.append({"issues": issues})        if not issues:            break        fixes = "\n".join(f"- {i}" for i in issues)        output = llm_fn(f"Task: {task}\n\nPrevious answer:\n{output}\n\n"                        f"Rewrite it, fixing every issue below and changing nothing else:\n{fixes}")    return {"output": output, "rounds": len(reviews), "reviews": reviews}

The tricky parts:

  • The external check runs first. If tests fail, their messages are the issues and no critic call is needed. The model's opinion is consulted only when the objective checks pass.
  • "Changing nothing else" in the revise prompt. Without it, the rewrite often fixes one issue and quietly breaks something that was right.
  • An unparseable critique ends the loop with the current output instead of crashing.
  • passed: false with an empty issues list is treated as passed — there is nothing concrete to fix.

Complexity: worst case is 1 generation + max_rounds × (critique + revision) = 1 + 2 × max_rounds model calls, or fewer when check replaces the critic. Space is the kept reviews, O(rounds).

A real-life example

A task with a format rule, a scripted model and a real external check (count the bullet points):

Python
task = "List 3 benefits of UPI as bullet points."replies = iter([    "- Instant\n- Free",                                   # first answer: only 2 bullets    "- Instant transfers\n- No fees for users\n- Works 24x7",  # revision    '{"passed": true, "issues": []}',                       # critic, after checks pass])def bullets(out: str) -> list[str]:    n = sum(line.startswith("- ") for line in out.splitlines())    return [] if n == 3 else [f"expected 3 bullets, found {n}"]result = reflect(task, lambda _: next(replies), check=bullets)print(result["rounds"], result["reviews"])# 2 [{'issues': ['expected 3 bullets, found 2']}, {'issues': []}]print(result["output"])# - Instant transfers# - No fees for users# - Works 24x7
roundexternal checkcritic called?action
1"expected 3 bullets, found 2"norevise with that issue
2passesyes → passedstop

Three model calls in total. The format problem was caught by 3 lines of Python, not by asking the model whether it had counted correctly.

Code-generation assistants use exactly this: generate, run the tests, feed failures back, and stop when the tests pass.

Follow-up questions to expect

  • "Why not just use a stronger model instead?" — Often that is the better trade: one call to a stronger model can beat three calls to a weaker one. Reflection earns its place when there is an external check to ground it.
  • "How do you know the critic is any good?" — Evaluate it on a labelled set of good and bad answers; measure how often it catches the bad ones and wrongly flags the good ones.
  • "Can reflection make answers worse?" — Yes. Revisions can over-edit or add hedging. Keep the round count low and compare quality with and without reflection.