Course Content
CrewAI Multi-Agents
9 sections · 53 lessons
What are the limitations of CrewAI?
What you need to know
A senior interviewer asks this to see whether you have run CrewAI in production. Name the limit, why it happens, and what you do about it.
| Limitation | Why it happens | Mitigation |
|---|---|---|
| Non-determinism in hierarchical mode | the manager LLM chooses order and wording at runtime | sequential by default; Flows for known branches |
| Cost and latency multiply | every agent is a loop of LLM calls | fewer agents, small models where possible, gates |
| Delegation loops | agents with allow_delegation=True pass work back and forth | delegation off for specialists, low max_iter |
| Framework-owned prompts | CrewAI builds part of every prompt | pin versions, golden-set tests before upgrades |
| No exactly-once | retries rerun tools | idempotency keys on side effects |
| Stale memory | memory retrieval can return an old fact | pass critical facts via context; scope or reset memory |
| Persona overhead | long backstories cost tokens | short, specific personas |
| Built-in test is limited | crewai test uses a generic judge (OpenAI only) | your own golden set and rubric |
What has improved
Some older criticisms are now partly out of date, and saying so shows you are current:
- Durability. Older answers said you could only
replayfrom a task. Current versions add checkpointing for crews, flows and agents (after every task by default, or on other events) and@persistfor Flow state. Recovery is still at event boundaries, not in the middle of a tool call. - Human in the loop. Flows now have
@human_feedback, including pausing and resuming later from a Slack or web approval. - Control flow. Flows give explicit routers and parallel branches, so you no longer need a hierarchical manager for branching.
Also worth knowing
CrewAI sends anonymous usage telemetry by default. For regulated work, turn it off with CREWAI_DISABLE_TELEMETRY=true (or OTEL_SDK_DISABLED=true), and leave share_crew off, since it shares task and agent text.
A real-life example
A bank piloted a hierarchical customer-support escalation crew with a manager and four specialists. In testing, the same complaint was routed to the card specialist on 7 runs and to the payments specialist on 3. One run had the investigator and writer delegate to each other 11 times before max_iter stopped them, costing about 12 times a normal run. A retried task also sent the customer two SMS messages.
The team kept CrewAI but changed the design: a Flow router picks the specialist crew with rules plus a small classifier, each crew is sequential, delegation is off, tools use idempotency keys, and telemetry is disabled. Routing became repeatable, and cost per ticket fell by about 60%.
Follow-up questions to expect
- "Is non-determinism a problem in sequential mode too?" — The order is fixed, but model outputs still vary. Schemas, guardrails and low temperature reduce it.
- "Would these limits make you pick another framework?" — Only if the workflow needs step-level durability or strict control everywhere. Otherwise the fixes above are enough.
- "What about vendor lock-in?" — The open-source library runs anywhere. The managed platform (CrewAI AMP) is optional; keep business logic in tools and schemas so it can move.