Course Content
Agents & Tools Interview Prep
6 sections · 40 lessons
How can you speed up slow agent workflows?
What you need to know
Find the time first
A trace for one run might look like this:
run 8f2c total 24.1 s model call 1 2.9 s (input 6k tokens) tool search_flights 3.8 s tool search_hotels 4.1 s <- ran after flights, not with it model call 2 3.4 s tool get_fare_rules 1.2 s tool get_fare_rules 1.1 s <- same arguments as before model call 3 4.9 s (input 31k tokens) model call 4 2.7 sThree problems are visible: serial tools, a repeated call, and a fat context in call 3.
Fixes mapped to causes
| Cause | Fix |
|---|---|
| Tools requested together but run one by one | Run them concurrently |
| Many small lookups | Merge into one tool that returns a summary |
| Same call repeated | Cache deterministic results per run |
| Slow external API | Timeout, retry with backoff, cache, or pre-fetch |
| Large context | Trim tool results, clear old ones, compact |
| Routine steps on a heavy setting | Lower effort or a smaller model for those routes |
| User stares at a blank screen | Stream text and show progress events |
Running tool calls concurrently
1import asyncio23async def run_tools(blocks):4 return await asyncio.gather(*(run_one(b) for b in blocks)) # not a for-loopIf the model returns three tool_use blocks, the turn now takes as long as the slowest tool, not the sum. Return all results in one message. Only make tools run one at a time when they conflict (two writes to the same record).
Why p95
Agent latency has a long tail: most runs take 8 seconds, a few take 60 because they wander. The average hides that tail; users remember it. Step caps and loop detection cut the tail more than any micro-optimisation.
A real-life example
A bank's operations assistant answered "Why did this customer's NEFT transfer fail?" in a median of 18 seconds, p95 of 55.
The trace showed the model calling get_customer, get_account and get_transfer in one turn — but the executor ran them in sequence (6 s). Then it called get_transfer_events 4 times with page numbers (5 s). And 1 in 12 runs looped on a transfer ID typo.
Fixes: concurrent execution (6 s became 2.4 s); a new get_transfer_timeline(transfer_id) tool that returns a merged, 20-line event summary (5 s became 1 s, and 3 model calls disappeared); a validator on the transfer ID format that returns a clear error; and progress messages streamed to the UI. Median fell to 7 seconds, p95 to 16.
Follow-up questions to expect
- "Can you parallelise model calls, not just tools?" — Yes, for independent sub-tasks: several sub-agents or several candidate answers at once. The run waits for the slowest.
- "Does streaming reduce actual latency?" — No, it reduces perceived latency. Users tolerate a 20-second agent that shows progress better than a silent 8-second one.
- "Would a faster model fix it?" — Sometimes, but first remove wasted steps and serial waits — they are free wins.