Course Content
CrewAI Multi-Agents
9 sections · 53 lessons
How do you optimize execution time in multi-agent systems?
What you need to know
If task B needs A's output, B waits for A. If tasks C and D need nothing from each other, they can run at the same time, and together they take only as long as the slower one.
Ways to shorten the path
- Parallel tasks in a crew. Mark independent tasks
async_execution=True; a later task that lists them incontextwaits for all of them.
1news = Task(description="Collect top business news for {date}", agent=news_agent,2 expected_output="5 items with sources", async_execution=True)3markets = Task(description="Collect index and currency closes for {date}",4 agent=markets_agent, expected_output="Figures with sources",5 async_execution=True)6digest = Task(description="Write the digest", agent=editor,7 expected_output="400 words", context=[news, markets])- Parallel Flow branches. Several listeners of the same method, or several
@start()methods, run together;@listen(and_(a, b, c))joins them. - Merge chatty tasks. Two tasks that always run together cost two full round trips of an agent loop. One task costs one.
- Faster models on the critical path. Use the big model on a branch that runs in parallel with something slower anyway.
- Caps.
max_iterandmax_execution_time(seconds) on agents stop one agent from adding minutes. - Fix slow tools. A 6-second API call in a 6-step loop is 36 seconds of waiting. Add timeouts, fetch in parallel inside the tool, and cache.
Many inputs at once
kickoff_for_each(inputs=[...])runs the crew once per input, one after another — convenient, but not faster.await crew.akickoff_for_each(inputs=[...])runs the inputs concurrently with native async, which is the one to use for speed.await crew.akickoff(...)is native async, good for many concurrent runs in a web service;kickoff_asyncwraps the sync version in a thread.
Feels faster vs is faster
stream=True on the crew streams output as it is written. Total time is the same, but users see progress in a second or two instead of a blank screen.
The limit you will hit
Running 10 agents in parallel can hit the provider's requests-per-minute limit, and then calls queue or fail. Set max_rpm on the crew and measure wall-clock time, not just "we made it parallel".
A real-life example
A daily news-digest flow must publish by 6:00 am. Version 1 ran six section writers one after another, about 40 seconds each, plus fact gathering and editing: about 5 minutes.
Version 2 gathers facts first (50 seconds), then runs the six section writers as parallel Flow listeners, joined with and_ before the editor. The six writers finish in about 55 seconds together (the slowest one), and the total drops to about 2 minutes. The first parallel attempt hit the provider's rate limit and two sections failed; setting max_rpm and adding retries in the LLM client fixed it.
Follow-up questions to expect
- "Does hierarchical process make things faster?" — Usually slower: the manager adds calls before, between and after workers.
- "What does
async_executiondo with a sync task after it?" — The next synchronous task that depends on them throughcontextwaits for them to finish. - "How do you find the slow step?" — A trace with timings per LLM call and tool call. Often the slow part is a tool, not the model.