Agents & Tools Interview Prep

Course Content

Agents & Tools Interview Prep

6 sections · 40 lessons

How can you speed up slow agent workflows?


Reading one slow run's tracemodel call1 — 2.9 ssearch_flights— 3.8 ssearch_hotels— 4.1 sget_fare_rules— 1.2 sget_fare_rules— 1.1 smodel call 3 —4.9 s, 31k innullwaitedfor flightssame args againfat context
The model asked for both searches at once; a serial executor turned its parallel request into a sum.

What you need to know

Find the time first

A trace for one run might look like this:

Text
run 8f2c  total 24.1 s  model call 1          2.9 s   (input 6k tokens)  tool search_flights   3.8 s  tool search_hotels    4.1 s   <- ran after flights, not with it  model call 2          3.4 s  tool get_fare_rules   1.2 s  tool get_fare_rules   1.1 s   <- same arguments as before  model call 3          4.9 s   (input 31k tokens)  model call 4          2.7 s

Three problems are visible: serial tools, a repeated call, and a fat context in call 3.

Fixes mapped to causes

CauseFix
Tools requested together but run one by oneRun them concurrently
Many small lookupsMerge into one tool that returns a summary
Same call repeatedCache deterministic results per run
Slow external APITimeout, retry with backoff, cache, or pre-fetch
Large contextTrim tool results, clear old ones, compact
Routine steps on a heavy settingLower effort or a smaller model for those routes
User stares at a blank screenStream text and show progress events

Running tool calls concurrently

Python
import asyncioasync def run_tools(blocks):    return await asyncio.gather(*(run_one(b) for b in blocks))   # not a for-loop

If the model returns three tool_use blocks, the turn now takes as long as the slowest tool, not the sum. Return all results in one message. Only make tools run one at a time when they conflict (two writes to the same record).

Why p95

Agent latency has a long tail: most runs take 8 seconds, a few take 60 because they wander. The average hides that tail; users remember it. Step caps and loop detection cut the tail more than any micro-optimisation.

A real-life example

A bank's operations assistant answered "Why did this customer's NEFT transfer fail?" in a median of 18 seconds, p95 of 55.

The trace showed the model calling get_customer, get_account and get_transfer in one turn — but the executor ran them in sequence (6 s). Then it called get_transfer_events 4 times with page numbers (5 s). And 1 in 12 runs looped on a transfer ID typo.

Fixes: concurrent execution (6 s became 2.4 s); a new get_transfer_timeline(transfer_id) tool that returns a merged, 20-line event summary (5 s became 1 s, and 3 model calls disappeared); a validator on the transfer ID format that returns a clear error; and progress messages streamed to the UI. Median fell to 7 seconds, p95 to 16.

Follow-up questions to expect

  • "Can you parallelise model calls, not just tools?" — Yes, for independent sub-tasks: several sub-agents or several candidate answers at once. The run waits for the slowest.
  • "Does streaming reduce actual latency?" — No, it reduces perceived latency. Users tolerate a 20-second agent that shows progress better than a silent 8-second one.
  • "Would a faster model fix it?" — Sometimes, but first remove wasted steps and serial waits — they are free wins.