Course Content
Scenario-Based AI Engineering Questions
26 sections · 146 lessons
Your LLM bill hits $40K instead of $4K. Someone added a feature that retries failed requests 10 times with full context. How do you build cost guardrails into an LLM-powered product?
What you need to know
Why retries exploded the bill
Each retry re-sends the whole context, and you pay for input tokens every time. Ten retries of a 20,000-token context is 200,000 input tokens for one user action. Worse, many failures are deterministic — a parse error or a too-long prompt fails the same way every time — so retrying them buys nothing.
| Failure | Retry? | Why |
|---|---|---|
| 429 rate limit | Yes, with backoff, honour Retry-After | Temporary |
| 5xx or timeout | Yes, once or twice | Usually temporary |
| 400 bad request, context too long | No | Will fail again |
| Output fails JSON schema | At most once, with the error message | Retrying blindly repeats it |
Guardrails in the request path
- Per call —
max_tokensalways set; estimate input tokens and reject or trim above a ceiling. - Per retry — at most two, exponential backoff with jitter, trimmed context on the retry.
- Per tenant — a daily token quota in Redis; when used up, fall back to a cache or a smaller model.
- Global — a circuit breaker: if retries exceed 10% of calls, stop retrying and page someone.
- Spend alert — hourly spend over 3x the trailing median pages the on-call.
A per-tenant quota check:
1import redis, datetime23r = redis.Redis()4DAILY_LIMIT = 2_000_000 # tokens per tenant per day56def charge(tenant: str, tokens: int) -> bool:7 key = f"tok:{tenant}:{datetime.date.today():%Y%m%d}"8 used = r.incrby(key, tokens)9 if used == tokens:10 r.expire(key, 2 * 86400) # first write today sets the expiry11 return used <= DAILY_LIMIT1213# before each call:14# if not charge(tenant, estimated_input + max_tokens): return degraded_answer()Charging the estimate before the call means a runaway loop hits the quota, not your invoice.
Attribution
Every call carries feature, tenant, model and prompt_version, and emits input and output tokens as metrics. The dashboards you want are cost per feature per day and cost per active user. Without them you know the bill rose but not which change caused it.
Structural savings, after the incident
Prompt caching for the fixed system prompt and shared context (most major providers now discount cached input tokens), routing simple requests to a cheaper model, a cache for repeated questions, and trimming retrieved context to what the reranker says is useful.
A real-life example
Scenario, numbers made up. An HR-tech startup's monthly model spend is normally around $4,000. A developer adds "retry up to 10 times" to a resume-parsing feature. Resumes over 30 pages hit the context limit, fail with a 400, and are retried ten times each with the full text. The month ends at $40,000.
The fix takes two days: retries only on 429 and 5xx, two at most; long resumes are split before the call; per-tenant quotas; and an hourly spend alert. Three weeks later a new bug causes a loop in another feature. The hourly alert fires after 50 minutes at about $300 of extra spend, and the circuit breaker has already stopped the retries.
Follow-up questions to expect
- "Why alert on hourly rate instead of a monthly budget?" — A monthly budget tells you after the damage. An hourly anomaly alert catches a runaway in under an hour, while it is still cheap.
- "What happens to users when a quota runs out?" — They get a degraded path, such as a cached answer, a smaller model or a clear message, not an unexplained error.
- "How do you stop this in code review?" — A shared client library that owns retries,
max_tokensand budgets, so feature code cannot call the provider directly.