Agents & Tools Interview Prep

Course Content

Agents & Tools Interview Prep

6 sections · 40 lessons

How do you manage token usage in long-running agent workflows?


Input tokens per step, 2,000-token tool results5k13k23k33k43k01234step 1step 5step 10step 15step 20The 20-step run totals about 480,000 input tokens; 400-token results cut it to about 140,000.
Every step resends everything before it, so shrinking each tool result pays off on every later turn.

What you need to know

Why it grows so fast

Suppose the system prompt and tools are 3,000 tokens and each step adds a 2,000-token tool result. The input for step n is about 3,000 + 2,000 × n.

Text
step 1:  5,000 input tokensstep 10: 23,000step 20: 43,000total over 20 steps: about 480,000 input tokens

The last step alone is 43,000 tokens, but the run costs almost half a million. Cutting each tool result from 2,000 to 400 tokens drops the total to about 140,000.

The levers, in order of payoff

  1. Small tool results at the source. Return only needed fields, cap rows, paginate, summarise long documents inside the tool. A tool that returns 8,000 tokens of raw JSON is a design bug.
  2. Clear stale tool results. Once a result has been used, keep a short placeholder. Anthropic's context editing (beta) can do this automatically, clearing old tool results while keeping the conversation's structure.
  3. Compact near the limit. Summarise older history into one block. Server-side compaction (beta) does this for you; if you write your own, keep the original goal, decisions, open questions and IDs word for word — those are what bad summaries lose.
  4. Move bulk data out. Write large intermediate results to a file or scratchpad and pass a reference; or use code execution / programmatic tool calling so data is filtered in a sandbox and only the answer reaches the model.
  5. Cache the stable prefix. Order the request as tools, then system prompt, then messages, with nothing that changes (timestamps, random IDs) near the front. Cached input is billed at a small fraction of the normal rate — check usage.cache_read_input_tokens to confirm it is working.
  6. Defer rarely used tools. Fifty tool schemas can be 10,000+ tokens on every call. Tool search with defer_loading: true loads only the ones the model asks for.

Measure it

Python
def log_usage(run_id, step, resp):    u = resp.usage    metrics.record(run_id=run_id, step=step,                   input=u.input_tokens, output=u.output_tokens,                   cache_read=u.cache_read_input_tokens or 0,                   cache_write=u.cache_creation_input_tokens or 0)

In the Anthropic API, input_tokens counts only the uncached part; cache reads and writes are reported separately. Alert when average tokens per run starts climbing — it often means a tool's output grew after an upstream change.

A real-life example

A SQL analytics agent at a retail company answered questions like "Compare weekend sales by city for Q2". Runs averaged 14 steps and 610,000 input tokens.

The trace showed three culprits: describe_table returned the full DDL of 80 tables (15,000 tokens) each time; run_query returned up to 1,000 rows; and the system prompt began with the current timestamp, so caching never hit.

Fixes: describe_table takes a table name and returns only that table; run_query returns 50 rows plus a total count, with a note to aggregate in SQL; the timestamp moved to the end of the user message; old query results are cleared after 3 turns. Average input fell to 120,000 tokens per run, cache reads covered most of the remainder, and answer quality went up — the model was no longer searching through walls of rows.

Follow-up questions to expect

  • "Compaction or clearing — which first?" — Clearing old tool results first; it is lossless for the conversation's reasoning. Compact only when the conversation itself is too long.
  • "Why not just use a model with a 1M-token context?" — It fits, but you pay for every token on every turn, latency rises with input size, and models attend less well to details buried in huge contexts.
  • "How do you know what a tool result costs?" — Count tokens on samples of real results (the API has a token-counting endpoint) and set a size budget per tool.