Scenario-Based AI Engineering Questions

Course Content

Scenario-Based AI Engineering Questions

26 sections · 146 lessons

Scenario – 6: Slow Sequential Execution


What you need to know

The scenario: a crew takes three minutes to produce a report, and users give up.

Where the time goes

Each agent step is at least one full LLM call. A crew's total time is roughly the sum of every call on the critical path. Independent tasks run one after another in a sequential process, so three 40-second research tasks cost two minutes even though none needs the others.

Structure: run independent tasks together

Python
market = Task(description="Market size for {product}", expected_output="...",              agent=market_researcher, async_execution=True)competitors = Task(description="Top competitors for {product}", expected_output="...",                   agent=competitor_researcher, async_execution=True)regulation = Task(description="Regulations affecting {product}", expected_output="...",                  agent=policy_researcher, async_execution=True)synthesis = Task(description="Write the brief", expected_output="...",                 agent=writer, context=[market, competitors, regulation])   # waits for all three

The three research tasks now take as long as the slowest one. For more complex patterns — branching, loops, fan-out and fan-in — a CrewAI Flow makes the parallelism explicit.

Per-call cost

FixWhy it helps
Tier modelsExtraction, classification and formatting run fine on a small, fast model; keep the big model for reasoning
Tool caching (cache=True, the default for agents)Repeated tool calls with the same input are not re-run
Short backstories and tool descriptionsThey are re-sent on every step of every task
Lower max_iterAgents often spend steps confirming what they know
Fewer agentsEach agent adds full round trips; remove any without a one-sentence job

Perceived latency

Stream the final task's output to the user and show progress for earlier tasks ("researching competitors…"). It does not reduce total time, but users feel the wait far less.

  1. Trace — per-task duration and tokens.
  2. Parallelise — independent tasks first; usually the biggest win.
  3. Tier models — the second biggest.
  4. Trim — prompts, iterations and agents.
  5. Re-measure — before changing anything else.

A real-life example

Scenario, numbers made up. A product team's crew writes launch briefs in about 180 seconds: three research tasks (40–50 seconds each), an analysis task (30 seconds) and a writing task (25 seconds), all sequential, all on the largest model.

Running the research tasks in parallel cuts about 90 seconds. Moving the analysis task's extraction steps to a small model and trimming 1,200-word backstories to 150 words saves another 25 seconds. Total time falls to about 65 seconds, and with progress messages and streaming, users see the first words of the brief in under 45 seconds. Quality on their 20-brief review set is unchanged.

Follow-up questions to expect

  • "Does parallel execution cost more?" — The same number of calls, so roughly the same tokens; you may hit rate limits sooner, so set max_rpm.
  • "When should you not parallelise?" — When a task needs another's output; running them together just makes the second one guess.
  • "How do you know a smaller model is good enough?" — Run your regression set with each model per task and compare.