Course Content
Agents & Tools Interview Prep
6 sections · 40 lessons
What strategies help reduce cost and latency in agent systems?
What you need to know
Cost levers
| Lever | How it helps | Watch out for |
|---|---|---|
| Prompt caching | Tools and system prompt are identical every turn; cached reads are much cheaper | A timestamp or unsorted tool list at the front breaks it |
| Small tool outputs | Fewer tokens resent every turn | Do not cut fields the model needs |
| Model tiering | Small model for routing and extraction | Measure accuracy on a labelled set |
| Effort setting | Lower effort for routine routes means fewer, merged tool calls | Hard tasks need higher effort |
| Fewer round trips | One get_order_summary instead of three lookups | Keep tools single-purpose enough to describe clearly |
| Batch API | Offline jobs (evals, backfills) at about half price | Results arrive within hours, not seconds |
Latency levers
- Parallel tool calls. When the model asks for three independent lookups in one turn, run them at the same time.
- Streaming. Users feel time-to-first-token, not total time. Stream text and show progress ("Searching flights…").
- Pre-fetch. If 90% of requests need the user's profile, load it before the first model call instead of waiting for the model to ask.
- Cache tool results. Deterministic lookups (fare rules, schema, FAQ) can be cached for minutes.
- Caps. Step limits stop a few runaway runs from dominating p95.
Per task, not per call
A cheaper model that needs 9 steps can cost more than a stronger model that finishes in 4. Compare configurations on cost per successful task and p95 latency per task, on the same evaluation set.
A real-life example
A travel assistant cost about ₹14 per conversation and took a median of 21 seconds. The trace showed:
- 6 model calls per conversation on average, each resending 9,000 tokens of tool definitions and instructions — and the cache hit rate was zero because the system prompt included "Current time: 14:32:07".
search_flightsandsearch_hotelswere requested together but run one after the other.- 3 of the 6 calls were simple look-ups (baggage rules, visa pages) done by the most expensive model at high effort.
Fixes: the time moved to the user message (cache hits rose to 85% of input tokens); tool calls in one turn run with asyncio.gather; baggage and visa look-ups became a pre-fetched context block; routine turns run at low effort. Result: about ₹5 per conversation and a median of 9 seconds, with the same task success rate on 500 test conversations.
Follow-up questions to expect
- "Where does prompt caching fail silently?" — Anything that changes the prefix: timestamps, per-request IDs, tools listed in a different order, or switching models mid-conversation (caches are per model).
- "Is a smaller model always cheaper?" — Per token, yes. Per completed task, not always — more steps and retries can erase the saving.
- "How do you cut latency without cutting quality?" — Parallelism, streaming and pre-fetching do not change what the model decides, so start there.