Scenario-Based AI Engineering Questions

Course Content

Scenario-Based AI Engineering Questions

26 sections · 146 lessons

Scenario – 5: Review Gate Workflow


What you need to know

The scenario: a crew drafts customer-facing content, and the business wants a review step before anything is published or sent.

Agent reviewer: structured, with criteria

Ask an LLM "is this good?" and it usually says yes. A useful reviewer checks against explicit criteria and the source material, and returns a verdict code can act on.

Python
from typing import Literalfrom pydantic import BaseModelclass Review(BaseModel):    verdict: Literal["pass", "revise", "reject"]    issues: list[str]    severity: Literal["low", "medium", "high"]review = Task(    description=("Check the draft against: 1) every price matches the rate card, "                 "2) no promise of delivery dates, 3) tone rules in the style guide. "                 "Cite the criterion for each issue."),    expected_output="A Review verdict",    output_pydantic=Review, agent=reviewer, context=[draft_task, rate_card_task],)

In a CrewAI Flow, a router reads verdict and sends the draft to a revise step or onward. Cap revision at two rounds; writer-reviewer ping-pong converges slowly and burns tokens.

Human gate: for real consequences

For anything sent to customers, spending money or changing records, add a human step: human_input=True on a task for simple cases, or a Flow that saves its state and pauses until a reviewer responds. Show the exact payload, and allow edits, not just approve or reject.

Keep the reviewer honest

  1. Sample — have humans review a weekly sample of drafts the agent passed and flagged.
  2. Measure — reviewer precision and recall against those human judgements.
  3. Log overrides — every case where a human disagreed is prime evaluation data.
  4. Narrow the human gate — auto-approve categories where agent and human agree almost always.

A real-life example

Scenario, numbers made up. An insurance company's crew drafts renewal emails. Its agent reviewer approves 99% of drafts. A human sample of 300 finds 11% contain a wrong premium or a promise the policy does not make — the reviewer caught almost none of them.

The team rewrites the reviewer with a structured verdict, explicit criteria and the rate card as context, caps revision at two rounds, and puts a human gate on every email above a premium threshold. Reviewer recall on the weekly sample rises from about 10% to 85%. After two months, simple renewals with no price change show 99% agent-human agreement and move to auto-approval; complex ones stay with humans.

Follow-up questions to expect

  • "Should the reviewer use the same model as the writer?" — A different model or at least a different prompt and criteria reduces the chance it shares the writer's blind spots.
  • "What if revision never passes?" — After two rounds, escalate to a human with the reviewer's issues attached.
  • "How do you stop humans rubber-stamping too?" — Show clear diffs and issues, keep queues small, and audit approvals occasionally.