Agents & Tools Interview Prep

Course Content

Agents & Tools Interview Prep

6 sections · 40 lessons

How should an agent handle conflicting outputs from multiple tools?


18,240 or 19,015 orders delivered in Pune?Comparedefinitionsand windowsNightlytable stopsat 11:30 pmNot a realconflict —use live tableReal conflict?Rank sources,then freshnessReport bothvalues and sourcesA later partner-API gap exposed a sync failure the same day.
Most disagreements are different questions, not different answers — and a surfaced real conflict is a bug found early.

What you need to know

Step 1: is it a real conflict?

Looks like a conflictActual reasonCheck
412 vs 408 active users7-day vs 30-day windowCompare definitions
₹48,200 vs ₹45,700 balanceOne includes a pending debitCompare "available" vs "ledger"
Flight at 06:10 vs 06:40Schedule changed; one source is cachedCompare as_of times
$120 vs ₹10,000Different currenciesNormalise units

This is why every tool should return metadata: as_of, source, currency, definition or filter used.

Step 2: resolve real conflicts by policy

  1. Source ranking — declared in code and in the prompt: primary system over replica, ledger over cache, first-party API over scraped page.
  2. Freshness — prefer the newer value when sources have equal authority.
  3. Tie-break — one extra call to a third source, at most.
  4. Escalate — for money, health, legal or safety numbers, show both values to a human.

Step 3: tell the user

A good answer names both values and why one was chosen: "Your available balance is ₹45,700 (core banking, 10:42). The monthly statement shows ₹48,200 as of 1 Sept; the difference is a pending ₹2,500 card payment."

In evaluations

Include test cases with planted conflicts. Count "silently chose one value" as a failure even if it chose the right one — next time it may not.

A real-life example

A SQL analytics agent is asked, "How many orders did we deliver in Pune yesterday?" It queries two tables: orders_summary (a nightly aggregate) says 18,240; orders (live table) says 19,015.

Checking definitions, it finds orders_summary excludes orders delivered after 11:30 pm because the nightly job runs then. Not a real conflict — a time-window difference. The agent answers 19,015 from the live table and explains the gap.

A week later, the same question shows 19,015 in orders and 21,300 in the logistics partner's API. Both claim full-day coverage. Policy says the internal orders table is the source of truth for "delivered" status. The agent reports 19,015, notes the partner count is 2,285 higher, and suggests checking for orders marked delivered by the partner but not yet synced. The data team finds a sync failure the same day — a conflict surfaced is a bug found.

Follow-up questions to expect

  • "Can the model decide which source to trust?" — It can apply a ranking you give it, but the ranking itself is a product decision, and for critical values it belongs in code.
  • "What if all sources are equally trusted?" — Prefer the freshest; if still unresolved, report both and escalate rather than averaging.
  • "How do you prevent conflicts in the first place?" — Fewer overlapping tools, clear definitions in tool descriptions, and one tool per business metric.