Course Content
Scenario-Based AI Engineering Questions
26 sections · 146 lessons
Scenario – 6: Slow Sequential Execution
What you need to know
The scenario: a crew takes three minutes to produce a report, and users give up.
Where the time goes
Each agent step is at least one full LLM call. A crew's total time is roughly the sum of every call on the critical path. Independent tasks run one after another in a sequential process, so three 40-second research tasks cost two minutes even though none needs the others.
Structure: run independent tasks together
1market = Task(description="Market size for {product}", expected_output="...",2 agent=market_researcher, async_execution=True)3competitors = Task(description="Top competitors for {product}", expected_output="...",4 agent=competitor_researcher, async_execution=True)5regulation = Task(description="Regulations affecting {product}", expected_output="...",6 agent=policy_researcher, async_execution=True)7synthesis = Task(description="Write the brief", expected_output="...",8 agent=writer, context=[market, competitors, regulation]) # waits for all threeThe three research tasks now take as long as the slowest one. For more complex patterns — branching, loops, fan-out and fan-in — a CrewAI Flow makes the parallelism explicit.
Per-call cost
| Fix | Why it helps |
|---|---|
| Tier models | Extraction, classification and formatting run fine on a small, fast model; keep the big model for reasoning |
Tool caching (cache=True, the default for agents) | Repeated tool calls with the same input are not re-run |
| Short backstories and tool descriptions | They are re-sent on every step of every task |
Lower max_iter | Agents often spend steps confirming what they know |
| Fewer agents | Each agent adds full round trips; remove any without a one-sentence job |
Perceived latency
Stream the final task's output to the user and show progress for earlier tasks ("researching competitors…"). It does not reduce total time, but users feel the wait far less.
- Trace — per-task duration and tokens.
- Parallelise — independent tasks first; usually the biggest win.
- Tier models — the second biggest.
- Trim — prompts, iterations and agents.
- Re-measure — before changing anything else.
A real-life example
Scenario, numbers made up. A product team's crew writes launch briefs in about 180 seconds: three research tasks (40–50 seconds each), an analysis task (30 seconds) and a writing task (25 seconds), all sequential, all on the largest model.
Running the research tasks in parallel cuts about 90 seconds. Moving the analysis task's extraction steps to a small model and trimming 1,200-word backstories to 150 words saves another 25 seconds. Total time falls to about 65 seconds, and with progress messages and streaming, users see the first words of the brief in under 45 seconds. Quality on their 20-brief review set is unchanged.
Follow-up questions to expect
- "Does parallel execution cost more?" — The same number of calls, so roughly the same tokens; you may hit rate limits sooner, so set
max_rpm. - "When should you not parallelise?" — When a task needs another's output; running them together just makes the second one guess.
- "How do you know a smaller model is good enough?" — Run your regression set with each model per task and compare.