Agents & Tools Interview Prep

Course Content

Agents & Tools Interview Prep

6 sections · 40 lessons

What is agent reflection, and how does it enhance performance?


What you need to know

Two kinds

Grounded reflection

  • Critique uses an external check
  • Tests, validators, query results, compilers
  • The critique contains new information
  • Reliable improvement

Self-review only

  • Model re-reads its own answer
  • No new evidence
  • Can confirm or "fix" wrongly
  • Small or no gain; extra cost

Grounded reflection with tools

Python
def answer_with_check(question, max_rounds=3):    sql = agent.write_sql(question)    for _ in range(max_rounds):        result = tools.run_readonly_query(sql, max_rows=50)        problems = checks.inspect(result)     # error? zero rows? nulls? totals don't match?        if not problems:            return agent.explain(question, sql, result)        sql = agent.revise_sql(question, sql, problems)   # critique is concrete    return escalate(question, sql, problems)

Each round costs a model call and a query, so it is capped. The critique is specific ("zero rows — city = 'bangalore' but values are 'Bengaluru'"), which is why the revision works.

Cheap checks that make good critics

  • Schema validation of a JSON output.
  • Running tests or a type-checker on generated code.
  • A SQL query's error, row count or a sanity total.
  • A second tool that cross-checks a number (for example, a ledger total against a report).
  • A separate critic prompt, ideally with a rubric, for text quality — weaker than a real check, but better than "check your work".

Rules

  • Stop early when the check passes.
  • Bound rounds — 2 or 3.
  • Watch for oscillation — if round 3 reverts to round 1's answer, stop and escalate.

A real-life example

A GitHub triage bot proposes one-line fixes for simple bugs. Version 1 wrote a patch and asked itself, "Is this patch correct?" It said yes about 95% of the time, but only 58% of patches passed CI.

Version 2 runs the repository's tests in a sandbox after writing the patch. If a test fails, the failure output goes back to the agent, and it gets two more attempts. Patches that pass rise to 81%, and the ones that still fail are labelled "needs human" instead of being posted with false confidence.

The cost: on average 1.6 extra model calls per issue and 90 seconds of CI time. For the maintainers, one fewer broken patch was worth far more.

Follow-up questions to expect

  • "Is reflection the same as self-consistency?" — No. Self-consistency samples several independent answers and votes; reflection revises one answer using a critique.
  • "Should the critic be a different model?" — It can help — a stronger or differently prompted critic avoids agreeing with itself — but an external check helps far more.
  • "When would you skip reflection?" — When there is no cheap objective check and latency matters, such as a chat reply; improve the first pass instead.