Scenario-Based AI Engineering Questions

Course Content

Scenario-Based AI Engineering Questions

26 sections · 146 lessons

Scenario – 10: Production Deployment Readiness


A crew run as a background jobAPI checksthe kill switchEnqueue job,return a job IDWorker runs crewunder a cost capClient pollsstatus or streamsStored resultwith its run IDOn relaunch day 1,100 briefs ran and the app stayed responsive.
Running a two-minute crew inside a web request lets one popular feature hold every worker and take the whole app down.

What you need to know

The scenario: a crew works in demos, and the team wants to put it in front of customers next month.

The gates

GateWhat "ready" means
Correctness30–50 real inputs in CI with checks on schema and key facts; variance measured over repeated runs; typed handoffs with guardrails
Reliabilitymax_iter and time limits per agent, bounded tool retries with backoff, a global run timeout, partial results instead of failures, idempotent writes
SafetyApproval for destructive actions, dry-run mode in CI, validated tool arguments, output moderation, tool output treated as untrusted
CostA hard per-run token and money cap that stops the run; model tiering; caching; alerts on cost per run
Observabilityrun_id, per-task spans, stored task outputs, dashboards for success rate, latency and cost per task
OperationsPinned model versions; prompts and crew definitions in git; canary rollout; a feature-flag kill switch; a written rollback plan

Run crews as background jobs

A crew run takes tens of seconds to minutes. Running it inside a web request ties up a server worker for the whole time; a traffic spike then exhausts all workers and the whole app stops responding.

Python
@app.post("/reports")def create_report(req: ReportRequest, user=Depends(auth)):    if not flags.enabled("report_crew"):                      # kill switch        raise HTTPException(503, "Report generation is paused")    job = queue.enqueue(run_report_crew, req.model_dump(), user.id,                        job_timeout=600, result_ttl=86400)    return {"job_id": job.id, "status_url": f"/reports/{job.id}"}def run_report_crew(payload, user_id):    budget = RunBudget(max_usd=2.00, max_tool_calls=60)       # the run stops when exceeded    crew = build_crew(budget=budget, run_id=current_job_id())    return crew.kickoff(inputs=payload).pydantic.model_dump()

Launch order

  1. Before launch — regression suite, kill switch, per-run cost cap, background job runner.
  2. Week one — dashboards per task, cost alerts, canary to a small group of users.
  3. Week two — dry-run CI for write tools, approval flows for any new actions.
  4. Ongoing — weekly review of failures and overrides; update the regression suite.

A real-life example

Scenario, numbers made up. A B2B startup launches a "market brief" crew inside its API server. On launch day, 200 users click "generate" within an hour; each run holds a worker for about 90 seconds, all workers fill up, and even the login page times out. One brief loops on a failing tool and costs $38.

The relaunch moves the crew to a queue with 10 workers and a status endpoint, adds a $2 per-run cap, a kill switch and a 40-input regression suite. On the second launch day, 1,100 briefs are generated; p95 wait including queueing is 3 minutes, the app stays responsive, the cost cap stops 7 runs, and the average cost per brief is $0.41.

Follow-up questions to expect

  • "What is the single most important gate?" — The per-run cost cap and kill switch together: they limit the damage of every failure you did not predict.
  • "How do you canary a crew?" — Route a small share of users to the new crew version, compare regression scores, cost and failure rates, then widen.
  • "How do you handle model upgrades?" — Pin versions, run the regression suite on the new model, and roll it out like any other change.