Scenario-Based AI Engineering Questions

Course Content

Scenario-Based AI Engineering Questions

26 sections · 146 lessons

Your LLM bill hits $40K instead of $4K. Someone added a feature that retries failed requests 10 times with full context. How do you build cost guardrails into an LLM-powered product?


How one feature multiplied the bill30-page resumeexceeds the contextProvider returns400, not a 429Code retries 10times, full textEvery retrypays for allinput tokensFix: retry only 429 and 5xx, twice at most, plus an hourly spend alert.
Retrying a deterministic failure buys nothing, and without a spend alert the damage is only visible on the invoice.

What you need to know

Why retries exploded the bill

Each retry re-sends the whole context, and you pay for input tokens every time. Ten retries of a 20,000-token context is 200,000 input tokens for one user action. Worse, many failures are deterministic — a parse error or a too-long prompt fails the same way every time — so retrying them buys nothing.

FailureRetry?Why
429 rate limitYes, with backoff, honour Retry-AfterTemporary
5xx or timeoutYes, once or twiceUsually temporary
400 bad request, context too longNoWill fail again
Output fails JSON schemaAt most once, with the error messageRetrying blindly repeats it

Guardrails in the request path

  1. Per call — max_tokens always set; estimate input tokens and reject or trim above a ceiling.
  2. Per retry — at most two, exponential backoff with jitter, trimmed context on the retry.
  3. Per tenant — a daily token quota in Redis; when used up, fall back to a cache or a smaller model.
  4. Global — a circuit breaker: if retries exceed 10% of calls, stop retrying and page someone.
  5. Spend alert — hourly spend over 3x the trailing median pages the on-call.

A per-tenant quota check:

Python
import redis, datetimer = redis.Redis()DAILY_LIMIT = 2_000_000                     # tokens per tenant per daydef charge(tenant: str, tokens: int) -> bool:    key = f"tok:{tenant}:{datetime.date.today():%Y%m%d}"    used = r.incrby(key, tokens)    if used == tokens:        r.expire(key, 2 * 86400)            # first write today sets the expiry    return used <= DAILY_LIMIT# before each call:# if not charge(tenant, estimated_input + max_tokens): return degraded_answer()

Charging the estimate before the call means a runaway loop hits the quota, not your invoice.

Attribution

Every call carries feature, tenant, model and prompt_version, and emits input and output tokens as metrics. The dashboards you want are cost per feature per day and cost per active user. Without them you know the bill rose but not which change caused it.

Structural savings, after the incident

Prompt caching for the fixed system prompt and shared context (most major providers now discount cached input tokens), routing simple requests to a cheaper model, a cache for repeated questions, and trimming retrieved context to what the reranker says is useful.

A real-life example

Scenario, numbers made up. An HR-tech startup's monthly model spend is normally around $4,000. A developer adds "retry up to 10 times" to a resume-parsing feature. Resumes over 30 pages hit the context limit, fail with a 400, and are retried ten times each with the full text. The month ends at $40,000.

The fix takes two days: retries only on 429 and 5xx, two at most; long resumes are split before the call; per-tenant quotas; and an hourly spend alert. Three weeks later a new bug causes a loop in another feature. The hourly alert fires after 50 minutes at about $300 of extra spend, and the circuit breaker has already stopped the retries.

Follow-up questions to expect

  • "Why alert on hourly rate instead of a monthly budget?" — A monthly budget tells you after the damage. An hourly anomaly alert catches a runaway in under an hour, while it is still cheap.
  • "What happens to users when a quota runs out?" — They get a degraded path, such as a cached answer, a smaller model or a clear message, not an unexplained error.
  • "How do you stop this in code review?" — A shared client library that owns retries, max_tokens and budgets, so feature code cannot call the provider directly.