Course Content
Agents & Tools Interview Prep
6 sections · 40 lessons
How do you manage token usage in long-running agent workflows?
What you need to know
Why it grows so fast
Suppose the system prompt and tools are 3,000 tokens and each step adds a 2,000-token tool result. The input for step n is about 3,000 + 2,000 × n.
step 1: 5,000 input tokensstep 10: 23,000step 20: 43,000total over 20 steps: about 480,000 input tokensThe last step alone is 43,000 tokens, but the run costs almost half a million. Cutting each tool result from 2,000 to 400 tokens drops the total to about 140,000.
The levers, in order of payoff
- Small tool results at the source. Return only needed fields, cap rows, paginate, summarise long documents inside the tool. A tool that returns 8,000 tokens of raw JSON is a design bug.
- Clear stale tool results. Once a result has been used, keep a short placeholder. Anthropic's context editing (beta) can do this automatically, clearing old tool results while keeping the conversation's structure.
- Compact near the limit. Summarise older history into one block. Server-side compaction (beta) does this for you; if you write your own, keep the original goal, decisions, open questions and IDs word for word — those are what bad summaries lose.
- Move bulk data out. Write large intermediate results to a file or scratchpad and pass a reference; or use code execution / programmatic tool calling so data is filtered in a sandbox and only the answer reaches the model.
- Cache the stable prefix. Order the request as tools, then system prompt, then messages, with nothing that changes (timestamps, random IDs) near the front. Cached input is billed at a small fraction of the normal rate — check
usage.cache_read_input_tokensto confirm it is working. - Defer rarely used tools. Fifty tool schemas can be 10,000+ tokens on every call. Tool search with
defer_loading: trueloads only the ones the model asks for.
Measure it
1def log_usage(run_id, step, resp):2 u = resp.usage3 metrics.record(run_id=run_id, step=step,4 input=u.input_tokens, output=u.output_tokens,5 cache_read=u.cache_read_input_tokens or 0,6 cache_write=u.cache_creation_input_tokens or 0)In the Anthropic API, input_tokens counts only the uncached part; cache reads and writes are reported separately. Alert when average tokens per run starts climbing — it often means a tool's output grew after an upstream change.
A real-life example
A SQL analytics agent at a retail company answered questions like "Compare weekend sales by city for Q2". Runs averaged 14 steps and 610,000 input tokens.
The trace showed three culprits: describe_table returned the full DDL of 80 tables (15,000 tokens) each time; run_query returned up to 1,000 rows; and the system prompt began with the current timestamp, so caching never hit.
Fixes: describe_table takes a table name and returns only that table; run_query returns 50 rows plus a total count, with a note to aggregate in SQL; the timestamp moved to the end of the user message; old query results are cleared after 3 turns. Average input fell to 120,000 tokens per run, cache reads covered most of the remainder, and answer quality went up — the model was no longer searching through walls of rows.
Follow-up questions to expect
- "Compaction or clearing — which first?" — Clearing old tool results first; it is lossless for the conversation's reasoning. Compact only when the conversation itself is too long.
- "Why not just use a model with a 1M-token context?" — It fits, but you pay for every token on every turn, latency rises with input size, and models attend less well to details buried in huge contexts.
- "How do you know what a tool result costs?" — Count tokens on samples of real results (the API has a token-counting endpoint) and set a size budget per tool.