Scenario-Based AI Engineering Questions

Course Content

Scenario-Based AI Engineering Questions

26 sections · 146 lessons

Your product team wants fully autonomous AI agents. Legal and compliance teams want strict human approval everywhere. How do you balance autonomy, compliance, and user experience in enterprise AI systems?


What you need to know

Four tiers of action

TierExamplesControl
Read-only, internalSearch, summarise, draftFully autonomous, logged
Reversible writesCreate a draft, update a ticket, schedule internallyAutonomous with undo and notification; reviewed after
External or irreversibleEmail a customer, move money, delete data, change production configHuman approval with an exact preview
Regulated decisionsCredit, hiring, medicalModel assists; a named human decides and signs

Legal gets control where it matters; users do not see a permission dialog on every step.

What compliance actually needs

  • An immutable audit log — prompt, retrieved context, tool calls, approvals and who approved.
  • Policy in code, not in the prompt — the prompt can be talked around; code cannot.
  • Scoped credentials per action — the agent can refund one order, not access the payments admin.
  • Value limits — refund up to ₹2,000 automatically, escalate above.
YAML
actions:  refund.issue:    tier: approval            # start here    limit_inr: 2000           # above this, always a human    promote_when:      min_cases: 300      max_override_rate: 0.02    demote_when:      override_rate_7d_above: 0.05  ticket.update:    tier: autonomous_with_undo  loan.decision:    tier: human_decides       # never promoted

The policy is data that legal can read and sign off. Promotion needs 300 cases with overrides under 2%; a rise above 5% in a week moves the action back to approval automatically.

  1. List every action the agent can take and assign a tier.
  2. Build the audit log before widening autonomy.
  3. Enforce in code — tiers, limits and scoped credentials.
  4. Measure overrides per action type.
  5. Promote and demote by the agreed rule, reviewed with legal each quarter.

A real-life example

Scenario, numbers made up. A bank's support agent can answer questions, raise disputes and issue fee refunds. Product wants it all autonomous; compliance wants every step approved. The pilot with full approval has agents clicking "approve" 40 times an hour, and they begin approving without reading.

The team tiers the actions. Answers and dispute drafts run freely with logging. Fee refunds start in approval mode with a ₹2,000 limit. After 6 weeks and 1,100 refunds with a 1.2% override rate, refunds under ₹2,000 become autonomous. Approvals fall from 40 to 5 an hour, now only for real decisions. In the next audit, compliance pulls any refund's full trail in minutes.

Follow-up questions to expect

  • "Who owns the policy file?" — Jointly: product proposes, legal and risk approve, engineering enforces. Changes go through review like code.
  • "How do you avoid rubber-stamp approvals?" — Fewer, better approvals: show a clear diff, only for high-risk actions, and track how long reviewers spend.
  • "Do regulations require a human in the loop?" — For some decisions, yes: the EU AI Act, for example, requires human oversight for high-risk systems such as hiring and credit scoring. Tie each regulated action to the rule that covers it.