Course Content
Agents & Tools Interview Prep
6 sections · 40 lessons
How do you enforce budget limits for agent tasks?
What you need to know
Three levels
| Level | Limits | Enforced by |
|---|---|---|
| Per step | max_tokens, tool timeout, tool output size | API parameter, executor |
| Per run | Steps, tokens, money, wall-clock time | Budget object in the loop |
| Per tenant / day | Money or tokens per customer | Shared counter (e.g. Redis) |
A budget object
1# rates live in config: USD per million tokens, per model2RATES = {"claude-opus-5": (5.00, 25.00), "claude-haiku-4-5": (1.00, 5.00)}34class Budget:5 def __init__(self, usd, steps):6 self.start_usd = usd7 self.usd_left, self.steps_left = usd, steps89 def charge(self, model, u):10 r_in, r_out = RATES[model]11 cost = (u.input_tokens * r_in12 + (u.cache_creation_input_tokens or 0) * r_in * 1.25 # 5-minute cache write13 + (u.cache_read_input_tokens or 0) * r_in * 0.1 # cache read14 + u.output_tokens * r_out) / 1_000_00015 self.usd_left -= cost16 self.steps_left -= 11718 def state(self):19 if self.usd_left <= 0 or self.steps_left <= 0:20 return "stop"21 if self.usd_left < 0.2 * self.start_usd or self.steps_left <= 2:22 return "wrap_up"23 return "ok"Charge from the response's real usage, not an estimate. Check state() before each call. Put the rates in configuration so a price change is a config change.
Degrade, don't guillotine
When the state is wrap_up, add an instruction like "Budget nearly used. Do not start new searches. Give your best answer with what you have and list what is unconfirmed." When it is stop, return the partial result and a clear status. A hard kill at 100% throws away everything the run learned.
Some APIs now let the model see its own budget. Anthropic's task budgets (beta) give the model a token allowance for an agentic loop that it paces itself against — it is advisory, so you still keep your hard limits.
Product decisions
The right ceiling depends on value: a support reply might get ₹4, a code-migration run ₹1,500. Make budgets part of each route's configuration, and alert when many runs hit the ceiling — that is a sign of a quality problem, not just a cost one.
A real-life example
A GitHub triage bot for a large open-source organisation runs on every new issue. One night, a bot in another repository opened 3,000 near-identical issues in an hour. Each triage run was capped at 10 steps, but nobody had a daily cap: the bill for that night was 40 times normal.
The team added:
- Per run: $0.08 and 10 steps; at 80%, wrap up with a label and a short comment.
- Per repository per day: $25, tracked in Redis; after that, new issues just get a
needs-triagelabel and no model call. - Anomaly alert: more than 5× the usual issue rate pages the on-call maintainer.
The next spam wave cost $25 for that repository and triggered the alert within 20 minutes.
Follow-up questions to expect
- "Why not just set
max_tokenslow?" — It limits one response, not the number of turns. A loop of short responses can still run for hours. - "How do you budget tool costs, not just model costs?" — Give each paid tool a cost and charge it to the same budget: a paid search API, a SQL warehouse query, a sandbox minute.
- "What should the user see when a budget is hit?" — The partial answer, what is missing, and an option to continue (which may require a human to approve more budget).