CrewAI Multi-Agents

Course Content

CrewAI Multi-Agents

9 sections · 53 lessons

How do you optimize execution time in multi-agent systems?


What you need to know

If task B needs A's output, B waits for A. If tasks C and D need nothing from each other, they can run at the same time, and together they take only as long as the slower one.

Ways to shorten the path

  • Parallel tasks in a crew. Mark independent tasks async_execution=True; a later task that lists them in context waits for all of them.
Python
news = Task(description="Collect top business news for {date}", agent=news_agent,            expected_output="5 items with sources", async_execution=True)markets = Task(description="Collect index and currency closes for {date}",               agent=markets_agent, expected_output="Figures with sources",               async_execution=True)digest = Task(description="Write the digest", agent=editor,              expected_output="400 words", context=[news, markets])
  • Parallel Flow branches. Several listeners of the same method, or several @start() methods, run together; @listen(and_(a, b, c)) joins them.
  • Merge chatty tasks. Two tasks that always run together cost two full round trips of an agent loop. One task costs one.
  • Faster models on the critical path. Use the big model on a branch that runs in parallel with something slower anyway.
  • Caps. max_iter and max_execution_time (seconds) on agents stop one agent from adding minutes.
  • Fix slow tools. A 6-second API call in a 6-step loop is 36 seconds of waiting. Add timeouts, fetch in parallel inside the tool, and cache.

Many inputs at once

  • kickoff_for_each(inputs=[...]) runs the crew once per input, one after another — convenient, but not faster.
  • await crew.akickoff_for_each(inputs=[...]) runs the inputs concurrently with native async, which is the one to use for speed.
  • await crew.akickoff(...) is native async, good for many concurrent runs in a web service; kickoff_async wraps the sync version in a thread.

Feels faster vs is faster

stream=True on the crew streams output as it is written. Total time is the same, but users see progress in a second or two instead of a blank screen.

The limit you will hit

Running 10 agents in parallel can hit the provider's requests-per-minute limit, and then calls queue or fail. Set max_rpm on the crew and measure wall-clock time, not just "we made it parallel".

A real-life example

A daily news-digest flow must publish by 6:00 am. Version 1 ran six section writers one after another, about 40 seconds each, plus fact gathering and editing: about 5 minutes.

Version 2 gathers facts first (50 seconds), then runs the six section writers as parallel Flow listeners, joined with and_ before the editor. The six writers finish in about 55 seconds together (the slowest one), and the total drops to about 2 minutes. The first parallel attempt hit the provider's rate limit and two sections failed; setting max_rpm and adding retries in the LLM client fixed it.

Follow-up questions to expect

  • "Does hierarchical process make things faster?" — Usually slower: the manager adds calls before, between and after workers.
  • "What does async_execution do with a sync task after it?" — The next synchronous task that depends on them through context waits for them to finish.
  • "How do you find the slow step?" — A trace with timings per LLM call and tool call. Often the slow part is a tool, not the model.