Agents & Tools Interview Prep

Course Content

Agents & Tools Interview Prep

6 sections · 40 lessons

How do you enforce budget limits for agent tasks?


What you need to know

Three levels

LevelLimitsEnforced by
Per stepmax_tokens, tool timeout, tool output sizeAPI parameter, executor
Per runSteps, tokens, money, wall-clock timeBudget object in the loop
Per tenant / dayMoney or tokens per customerShared counter (e.g. Redis)

A budget object

Python
# rates live in config: USD per million tokens, per modelRATES = {"claude-opus-5": (5.00, 25.00), "claude-haiku-4-5": (1.00, 5.00)}class Budget:    def __init__(self, usd, steps):        self.start_usd = usd        self.usd_left, self.steps_left = usd, steps    def charge(self, model, u):        r_in, r_out = RATES[model]        cost = (u.input_tokens * r_in                + (u.cache_creation_input_tokens or 0) * r_in * 1.25   # 5-minute cache write                + (u.cache_read_input_tokens or 0) * r_in * 0.1        # cache read                + u.output_tokens * r_out) / 1_000_000        self.usd_left -= cost        self.steps_left -= 1    def state(self):        if self.usd_left <= 0 or self.steps_left <= 0:            return "stop"        if self.usd_left < 0.2 * self.start_usd or self.steps_left <= 2:            return "wrap_up"        return "ok"

Charge from the response's real usage, not an estimate. Check state() before each call. Put the rates in configuration so a price change is a config change.

Degrade, don't guillotine

When the state is wrap_up, add an instruction like "Budget nearly used. Do not start new searches. Give your best answer with what you have and list what is unconfirmed." When it is stop, return the partial result and a clear status. A hard kill at 100% throws away everything the run learned.

Some APIs now let the model see its own budget. Anthropic's task budgets (beta) give the model a token allowance for an agentic loop that it paces itself against — it is advisory, so you still keep your hard limits.

Product decisions

The right ceiling depends on value: a support reply might get ₹4, a code-migration run ₹1,500. Make budgets part of each route's configuration, and alert when many runs hit the ceiling — that is a sign of a quality problem, not just a cost one.

A real-life example

A GitHub triage bot for a large open-source organisation runs on every new issue. One night, a bot in another repository opened 3,000 near-identical issues in an hour. Each triage run was capped at 10 steps, but nobody had a daily cap: the bill for that night was 40 times normal.

The team added:

  • Per run: $0.08 and 10 steps; at 80%, wrap up with a label and a short comment.
  • Per repository per day: $25, tracked in Redis; after that, new issues just get a needs-triage label and no model call.
  • Anomaly alert: more than 5× the usual issue rate pages the on-call maintainer.

The next spam wave cost $25 for that repository and triggered the alert within 20 minutes.

Follow-up questions to expect

  • "Why not just set max_tokens low?" — It limits one response, not the number of turns. A loop of short responses can still run for hours.
  • "How do you budget tool costs, not just model costs?" — Give each paid tool a cost and charge it to the same budget: a paid search API, a SQL warehouse query, a sandbox minute.
  • "What should the user see when a budget is hit?" — The partial answer, what is missing, and an option to continue (which may require a human to approve more budget).