Course Content
Scenario-Based AI Engineering Questions
26 sections · 146 lessons
Scenario – 5: Review Gate Workflow
What you need to know
The scenario: a crew drafts customer-facing content, and the business wants a review step before anything is published or sent.
Agent reviewer: structured, with criteria
Ask an LLM "is this good?" and it usually says yes. A useful reviewer checks against explicit criteria and the source material, and returns a verdict code can act on.
1from typing import Literal2from pydantic import BaseModel34class Review(BaseModel):5 verdict: Literal["pass", "revise", "reject"]6 issues: list[str]7 severity: Literal["low", "medium", "high"]89review = Task(10 description=("Check the draft against: 1) every price matches the rate card, "11 "2) no promise of delivery dates, 3) tone rules in the style guide. "12 "Cite the criterion for each issue."),13 expected_output="A Review verdict",14 output_pydantic=Review, agent=reviewer, context=[draft_task, rate_card_task],15)In a CrewAI Flow, a router reads verdict and sends the draft to a revise step or onward. Cap revision at two rounds; writer-reviewer ping-pong converges slowly and burns tokens.
Human gate: for real consequences
For anything sent to customers, spending money or changing records, add a human step: human_input=True on a task for simple cases, or a Flow that saves its state and pauses until a reviewer responds. Show the exact payload, and allow edits, not just approve or reject.
Keep the reviewer honest
- Sample — have humans review a weekly sample of drafts the agent passed and flagged.
- Measure — reviewer precision and recall against those human judgements.
- Log overrides — every case where a human disagreed is prime evaluation data.
- Narrow the human gate — auto-approve categories where agent and human agree almost always.
A real-life example
Scenario, numbers made up. An insurance company's crew drafts renewal emails. Its agent reviewer approves 99% of drafts. A human sample of 300 finds 11% contain a wrong premium or a promise the policy does not make — the reviewer caught almost none of them.
The team rewrites the reviewer with a structured verdict, explicit criteria and the rate card as context, caps revision at two rounds, and puts a human gate on every email above a premium threshold. Reviewer recall on the weekly sample rises from about 10% to 85%. After two months, simple renewals with no price change show 99% agent-human agreement and move to auto-approval; complex ones stay with humans.
Follow-up questions to expect
- "Should the reviewer use the same model as the writer?" — A different model or at least a different prompt and criteria reduces the chance it shares the writer's blind spots.
- "What if revision never passes?" — After two rounds, escalate to a human with the reviewer's issues attached.
- "How do you stop humans rubber-stamping too?" — Show clear diffs and issues, keep queues small, and audit approvals occasionally.