Course Content
CrewAI Multi-Agents
9 sections · 53 lessons
How do you avoid unnecessary agent invocations?
What you need to know
Each agent invocation is a loop of LLM calls, often 3–10 of them. A crew that runs 5 agents for every input, including trivial ones, pays for all of them every time.
Four ways to skip work
1. Do not use an agent for deterministic work. If code can do it exactly — convert currency, format a date, look up an order by ID, fill a template — code is faster, free and testable. An agent is for open-ended judgement.
2. Gate the expensive path. A single cheap call or a rule decides whether the full crew runs.
1class SupportFlow(Flow[TicketState]):2 @start()3 def triage(self):4 out = triage_agent.kickoff(self.state.text, response_format=Triage)5 self.state.triage = out.pydantic67 @router(triage)8 def route(self):9 if self.state.triage.kind == "faq":10 return "faq"11 return "escalate"1213 @listen("faq")14 def answer_faq(self):15 return faq_answer(self.state.triage.topic) # template, no agent1617 @listen("escalate")18 def run_crew(self):19 return escalation_crew.kickoff(inputs={"ticket": self.state.text})3. Branch inside a crew. A ConditionalTask runs only if a function of the previous task's output returns True, for example "fetch more data only if fewer than 10 results came back".
4. Turn off delegation on specialists. allow_delegation is False by default now, but many examples set it to True. Each delegation starts another agent loop with its own context. Keep it on only for a manager or coordinator.
Caching at the right level
- Tool cache — opt-in with
Crew(cache=True); identical tool calls in one run then reuse the result. An agent's owncacheflag (defaultTrue) only means it takes part when the crew turns caching on. Do not cache live-data tools (order status, stock prices) or tools that change state, unless the tool'scache_functionblocks caching for those calls. - Your own result cache — if the same input arrives again (same claim ID, same topic today), return the stored result instead of rerunning the crew.
- Provider prompt caching — lowers the cost of calls you still make.
A real-life example
A telecom's customer-support escalation crew received every ticket, about 20,000 a day. Analysis showed 65% were simple: "how do I recharge", "what is my plan", "change my address".
The team added a triage step (a small model, about ₹0.10 per ticket) and a Flow router. FAQ tickets get a templated answer filled with account data from an API, with no agent. Only 7,000 tickets reach the four-agent crew. They also found the investigator had allow_delegation=True and delegated to the writer on 30% of tickets, adding an extra agent run each time; they turned it off.
Daily cost fell from about ₹1.8 lakh to ₹55,000, and median response time for simple tickets dropped from 40 seconds to 3.
Follow-up questions to expect
- "What if triage sends a hard ticket down the cheap path?" — Measure triage accuracy on labelled tickets, and add a way back: if the customer replies "that didn't help", route to the crew.
- "Is
ConditionalTaskor a Flow router better?" —ConditionalTaskfor a skip inside one crew; a Flow router when the branches are different crews or plain code. - "How do you find unnecessary calls?" — Look at traces for agents whose output is never used downstream, and for delegations and repeated tool calls.