Scenario-Based AI Engineering Questions

Course Content

Scenario-Based AI Engineering Questions

26 sections · 146 lessons

Scenario – 8: Human Approval Gate


A refund that waits for a personagentproposes refundpolicy table:needs approvalinterrupt,state checkpointedrevieweredits in SlackCommand(resume)hours laterrefundissued oncenullcode decidessurvivesdeploysidempotentThe node re-runs from the top on resume, so side effects go after the interrupt.
Rejection rates per action type are the evidence that later lets small, safe refunds skip the gate.

What you need to know

The scenario: an agent can issue refunds, send emails or change records. Some actions must wait for a human to approve, possibly for hours.

The code

Python
from langgraph.types import interrupt, CommandNEEDS_APPROVAL = {"issue_refund", "send_email", "delete_record"}     # policy lives in codedef approval(state):    action = state["proposed_action"]    if action["tool"] not in NEEDS_APPROVAL:        return {"approved": True, "payload": action}    decision = interrupt({"action": action, "sources": state["sources"]})    return {"approved": decision["ok"], "payload": decision.get("edited", action),            "reviewer": decision["reviewer"]}# later, from a Slack button or web handler:graph.invoke(Command(resume={"ok": True, "reviewer": "priya@company.in"}),             config={"configurable": {"thread_id": tid}})

When the node resumes, it runs again from the start and interrupt returns the resume value. So code before interrupt must be safe to run twice — never send the email before the pause.

Design points

  1. Gate by blast radius — writes, payments, external messages and deletes need approval; reads do not.
  2. Show the exact payload — the actual refund amount and recipient, plus the reasoning and sources.
  3. Allow edits — an edited payload is the most useful feedback you will collect.
  4. Expire and survive — pending approvals have a TTL, and a Postgres checkpointer keeps them through deploys.
  5. Audit — reviewer, time, original payload, edited payload and outcome.
MetricWhat it tells you
Time to approvalWhether the gate is slowing customers down
Rejection rate per action typeWhere the agent is still wrong
Edit distance on approved payloadsHow much reviewers are fixing

When an action type's rejection and edit rates stay very low over many cases, that is evidence for auto-approving it, perhaps below an amount limit. Autonomy is earned with data.

A real-life example

Scenario, numbers made up. An e-commerce support agent proposes refunds. All refunds pause for approval in a Slack channel with the order, the amount, the customer's message and the policy clause the agent used. Reviewers can approve, reject or edit the amount.

After six weeks and about 9,000 refunds, refunds under ₹500 for "item not delivered" with a courier confirmation have a 0.4% rejection rate, while "damaged item" refunds have 11%. The team auto-approves the first type below ₹500 and keeps the gate for the rest. Median time to refund for small cases falls from 3 hours to 2 minutes, while reviewers focus on the risky cases.

Follow-up questions to expect

  • "What happens if the reviewer never responds?" — The approval expires, the user is told, and the case goes to a human queue; the graph should not wait forever.
  • "Why can't the model decide when it needs approval?" — A prompt-injected or confused model would decide it does not; the gate has to be a deterministic rule.
  • "What if the server restarts while waiting?" — With a durable checkpointer, the paused state is in the database, and the resume call works on any instance.