Course Content
Scenario-Based AI Engineering Questions
26 sections · 146 lessons
Scenario – 10: Production Deployment Readiness
What you need to know
The scenario: a crew works in demos, and the team wants to put it in front of customers next month.
The gates
| Gate | What "ready" means |
|---|---|
| Correctness | 30–50 real inputs in CI with checks on schema and key facts; variance measured over repeated runs; typed handoffs with guardrails |
| Reliability | max_iter and time limits per agent, bounded tool retries with backoff, a global run timeout, partial results instead of failures, idempotent writes |
| Safety | Approval for destructive actions, dry-run mode in CI, validated tool arguments, output moderation, tool output treated as untrusted |
| Cost | A hard per-run token and money cap that stops the run; model tiering; caching; alerts on cost per run |
| Observability | run_id, per-task spans, stored task outputs, dashboards for success rate, latency and cost per task |
| Operations | Pinned model versions; prompts and crew definitions in git; canary rollout; a feature-flag kill switch; a written rollback plan |
Run crews as background jobs
A crew run takes tens of seconds to minutes. Running it inside a web request ties up a server worker for the whole time; a traffic spike then exhausts all workers and the whole app stops responding.
1@app.post("/reports")2def create_report(req: ReportRequest, user=Depends(auth)):3 if not flags.enabled("report_crew"): # kill switch4 raise HTTPException(503, "Report generation is paused")5 job = queue.enqueue(run_report_crew, req.model_dump(), user.id,6 job_timeout=600, result_ttl=86400)7 return {"job_id": job.id, "status_url": f"/reports/{job.id}"}89def run_report_crew(payload, user_id):10 budget = RunBudget(max_usd=2.00, max_tool_calls=60) # the run stops when exceeded11 crew = build_crew(budget=budget, run_id=current_job_id())12 return crew.kickoff(inputs=payload).pydantic.model_dump()Launch order
- Before launch — regression suite, kill switch, per-run cost cap, background job runner.
- Week one — dashboards per task, cost alerts, canary to a small group of users.
- Week two — dry-run CI for write tools, approval flows for any new actions.
- Ongoing — weekly review of failures and overrides; update the regression suite.
A real-life example
Scenario, numbers made up. A B2B startup launches a "market brief" crew inside its API server. On launch day, 200 users click "generate" within an hour; each run holds a worker for about 90 seconds, all workers fill up, and even the login page times out. One brief loops on a failing tool and costs $38.
The relaunch moves the crew to a queue with 10 workers and a status endpoint, adds a $2 per-run cap, a kill switch and a 40-input regression suite. On the second launch day, 1,100 briefs are generated; p95 wait including queueing is 3 minutes, the app stays responsive, the cost cap stops 7 runs, and the average cost per brief is $0.41.
Follow-up questions to expect
- "What is the single most important gate?" — The per-run cost cap and kill switch together: they limit the damage of every failure you did not predict.
- "How do you canary a crew?" — Route a small share of users to the new crew version, compare regression scores, cost and failure rates, then widen.
- "How do you handle model upgrades?" — Pin versions, run the regression suite on the new model, and roll it out like any other change.