Agentic AI Patterns

Course Content

Agentic AI Patterns

9 sections · 50 lessons

How does an Agentic AI system decide between parallel vs. sequential execution?


Four independent lookups, in seconds1.20.82.01.50123flightshotelspolicy —the slowestvisa rulesIn sequence: 5.5 s. In parallel: 2.0 s, the slowest branch.
Parallelism buys time, not money — and the bookings that follow must still run one after another.

What you need to know

The latency maths

Four independent tool calls take 1.2 s, 0.8 s, 2.0 s and 1.5 s.

Text
Sequential: 1.2 + 0.8 + 2.0 + 1.5 = 5.5 sParallel:   max(1.2, 0.8, 2.0, 1.5) = 2.0 s

Parallel is bounded by the slowest branch. Cost is the same or higher (retries multiply), so parallelism buys time, not money.

Run in parallel when

  • Read-only lookups that do not depend on each other.
  • Research across several sources.
  • Per-file or per-record analysis.
  • Several samples for voting.

Run in sequence when

  • Step N needs step N−1's output ("find flight", then "hold that flight").
  • Side effects must be ordered ("create PO", then "send PO").
  • An early cheap check can stop the run (fraud score first, before an expensive review).
  • A human must approve between steps.

Models can request parallel calls

Most current models can return several tool calls in one turn. Your executor should run them concurrently if they are read-only, and one by one if any writes.

Python
import asyncioasync def gather_with_limits(calls, call_tool, timeout=2.5, max_parallel=4):    sem = asyncio.Semaphore(max_parallel)     # respect provider rate limits    async def one(name, args):        async with sem:            try:                return name, await asyncio.wait_for(call_tool(name, args), timeout)            except Exception as e:                return name, f"FAILED: {type(e).__name__}"    return dict(await asyncio.gather(*(one(n, a) for n, a in calls)))

Each branch gets its own timeout and its failure is returned as a value, so one slow service does not sink the whole step. The semaphore caps concurrency, because provider and API rate limits are usually the real ceiling. The merge is by tool name, not by finish order, so it is deterministic.

Partial failure policy

Decide it up front: fail fast, or continue with what returned and say what is missing. For a travel search, continuing is fine; for a payment reconciliation, it usually is not.

A real-life example

A travel-booking agent receives "Book Mumbai to Singapore for the 10th to 14th, within policy".

Phase 1, parallel (reads): search_flights, search_hotels, get_travel_policy, get_visa_rules. Timings 1.2 s, 0.8 s, 2.0 s, 1.5 s. Parallel: 2.0 s instead of 5.5 s. The visa service times out; the agent continues and tells the traveller "visa rules could not be checked; please confirm your visa status".

Phase 2, sequential (writes): hold the chosen flight, then hold the hotel for the matching dates, then ask for manager approval because the fare is above the cap, then confirm both. Holding the hotel before the flight is confirmed could leave a hotel booked with no flight, so order matters.

Whole run: about 9 seconds of tool time instead of about 15, and no risk of half-booked trips.

Follow-up questions to expect

  • "Does parallel execution save money?" — No. It saves time. Total tokens and API calls are the same, and retries across many branches can add cost.
  • "How do you handle a DAG of steps?" — Group steps into layers by dependency; run each layer in parallel and the layers in order.
  • "What about parallel writes?" — Avoid them unless the writes are independent and idempotent. Ordering bugs in writes are hard to undo.