Course Content
CrewAI Multi-Agents
9 sections · 53 lessons
How do you balance quality vs cost in CrewAI pipelines?
What you need to know
"Better quality" and "lower cost" are only useful words if you can measure them.
- Quality: a score on a golden set, from a rubric or checklist. For example, "percent of claims with the correct coverage decision".
- Cost: rupees or dollars per run, plus latency if users wait.
A ladder to climb, measuring at each step
- One agent, one clear task — sharp
expected_output, good description. - Tools — for facts the model cannot know, like policy text or order status.
- Schema and guardrails —
output_pydanticand code checks. Cheap, reliable quality. - A critic or review step — only where the golden set shows errors a reviewer would catch.
- A stronger model on one agent — the one where measurement shows it matters.
- More agents — only for a specific, measured gap: different tools, different permissions, or parallel work.
Steps 1–3 are cheap and often give most of the quality. Steps 4–6 multiply cost, so each needs evidence.
Compare setups as points
For each setup, record (cost per run, score). Keep only the setups where nothing else is both cheaper and better. From those, pick the cheapest one above your quality bar.
When high cost is right
Some decisions justify a costly setup: a wrong ₹10 lakh claim approval, a legal answer, a medical triage. Put the expensive path only where the stakes are, and route the rest to the cheap path.
A real-life example
An insurer tested five setups of its insurance-claims review crew on 100 labelled claims. Required quality: at least 92% correct decisions.
| Setup | Cost per claim | Correct |
|---|---|---|
| A. One agent, small model | ₹2 | 81% |
| B. A + policy tool + schema + guardrails | ₹3 | 91% |
| C. B + critic agent | ₹6 | 93% |
| D. B + strong model on policy check only | ₹5 | 94% |
| E. Four agents, strong model everywhere | ₹22 | 94% |
Setup D wins: it passes the bar at a quarter of E's cost. C also passes, but costs more for a lower score. The biggest jump (81% to 91%) came from the tool and schema, not from more agents. The team also sends claims above ₹5 lakh to setup E plus a human, because a wrong decision there costs far more than ₹22.
Follow-up questions to expect
- "How do you know the difference is real and not noise?" — Use enough cases (100 or more for small differences), run each setup more than once, and look at which cases changed, not just the total.
- "What about latency?" — Treat it as a third number. A critic step may pass on quality and cost but add 20 seconds a user will not accept.
- "How often do you re-check?" — When models, prompts or traffic change, and on a schedule, because new cheaper models may pass your bar.