Agents & Tools Interview Prep

Course Content

Agents & Tools Interview Prep

6 sections · 40 lessons

How does human-in-the-loop improve agent reliability?


A transfer the model proposedModel callscreate_transfer20,000Policy: over10,000, needs PINRun saved,card showspayee, bank, digitsCustomerapprovesin the appRun resumes fromthe checkpointAdding bank and last 4 digits to the card ended wrong-payee approvals.
Approval works only when the human sees the resolved action and answers through a channel the model cannot write to.

What you need to know

Where to put humans

PlacementTriggerExample
Approve actionSide-effect tool with real costTransfer ₹20,000; email 4,812 customers
ClarifyAmbiguous or missing input"Which of your two accounts?"
Review planLong or expensive runApprove a 12-step data migration plan once
EscalateRepeated failure or low confidenceHand the chat to a human agent after 2 failed attempts

What a good approval shows

Show the resolved action, not the intent:

  • Weak: "The assistant wants to make a transfer."
  • Strong: "Send ₹20,000 from Savings ••0042 to Priya Sharma (UPI priya@okbank). Balance after: ₹28,200."

A person can only catch a mistake they can see.

A durable pause

With LangGraph, interrupt() stops the graph and saves state through the checkpointer; the run resumes when you send the human's answer:

Python
from langgraph.types import interrupt, Commanddef confirm_transfer(state):    answer = interrupt({"action": "create_transfer",                        "payee": state["payee"], "amount_inr": state["amount"]})    return {"approved": answer == "approve"}# later, when the customer taps a button:graph.invoke(Command(resume="approve"), config={"configurable": {"thread_id": "chat-91"}})

Nothing is held in memory while waiting, so the customer can approve from another device an hour later. Without a framework, the same idea is: save the pending tool call and conversation to a database, return to the user, and resume the loop with the decision as the tool_result.

Avoiding approval fatigue

If every step needs approval, people click "Approve" without reading. Tier the rules:

  • Auto-approve reads and reversible low-value actions.
  • Confirm writes above a threshold or visible to others.
  • Two-person approval for rare, high-impact actions.

Track the approval rate. If humans approve 100% of a category for months, it may not need a human; if they reject 30%, the agent needs work.

A real-life example

A bank's assistant can look up balances freely but must confirm transfers. The rules:

  • get_balance, list_transactions — no approval.
  • create_transfer up to ₹10,000 to a saved payee — one tap in the app.
  • create_transfer to a new payee or above ₹10,000 — in-app confirmation plus the customer's UPI PIN, entered in the bank's own secure screen, never typed into the chat.
  • Adding a new payee — cooling-off period set by the bank's rules; the agent cannot shorten it.

In the first month, customers rejected 4% of proposed transfers — mostly because the agent chose the wrong "Priya" from the payee list. The team added the payee's bank and last 4 digits to the confirmation card, and rejections due to the wrong payee dropped to almost zero. The approval step had surfaced a real bug.

Follow-up questions to expect

  • "Doesn't human approval kill automation?" — Only if you gate everything. Gate the few actions where a mistake is expensive; automate the rest.
  • "How do you handle a human who never responds?" — Time-outs: cancel or expire the pending action after a set time, and notify the user.
  • "Can the agent approve itself?" — No. Approval must come from a channel the model cannot write to, such as a signed UI event.