Agents & Tools Interview Prep

Course Content

Agents & Tools Interview Prep

6 sections · 40 lessons

What strategies help reduce cost and latency in agent systems?


What you need to know

Cost levers

LeverHow it helpsWatch out for
Prompt cachingTools and system prompt are identical every turn; cached reads are much cheaperA timestamp or unsorted tool list at the front breaks it
Small tool outputsFewer tokens resent every turnDo not cut fields the model needs
Model tieringSmall model for routing and extractionMeasure accuracy on a labelled set
Effort settingLower effort for routine routes means fewer, merged tool callsHard tasks need higher effort
Fewer round tripsOne get_order_summary instead of three lookupsKeep tools single-purpose enough to describe clearly
Batch APIOffline jobs (evals, backfills) at about half priceResults arrive within hours, not seconds

Latency levers

  • Parallel tool calls. When the model asks for three independent lookups in one turn, run them at the same time.
  • Streaming. Users feel time-to-first-token, not total time. Stream text and show progress ("Searching flights…").
  • Pre-fetch. If 90% of requests need the user's profile, load it before the first model call instead of waiting for the model to ask.
  • Cache tool results. Deterministic lookups (fare rules, schema, FAQ) can be cached for minutes.
  • Caps. Step limits stop a few runaway runs from dominating p95.

Per task, not per call

A cheaper model that needs 9 steps can cost more than a stronger model that finishes in 4. Compare configurations on cost per successful task and p95 latency per task, on the same evaluation set.

A real-life example

A travel assistant cost about ₹14 per conversation and took a median of 21 seconds. The trace showed:

  • 6 model calls per conversation on average, each resending 9,000 tokens of tool definitions and instructions — and the cache hit rate was zero because the system prompt included "Current time: 14:32:07".
  • search_flights and search_hotels were requested together but run one after the other.
  • 3 of the 6 calls were simple look-ups (baggage rules, visa pages) done by the most expensive model at high effort.

Fixes: the time moved to the user message (cache hits rose to 85% of input tokens); tool calls in one turn run with asyncio.gather; baggage and visa look-ups became a pre-fetched context block; routine turns run at low effort. Result: about ₹5 per conversation and a median of 9 seconds, with the same task success rate on 500 test conversations.

Follow-up questions to expect

  • "Where does prompt caching fail silently?" — Anything that changes the prefix: timestamps, per-request IDs, tools listed in a different order, or switching models mid-conversation (caches are per model).
  • "Is a smaller model always cheaper?" — Per token, yes. Per completed task, not always — more steps and retries can erase the saving.
  • "How do you cut latency without cutting quality?" — Parallelism, streaming and pre-fetching do not change what the model decides, so start there.