Course Content
Scenario-Based AI Engineering Questions
26 sections · 146 lessons
Your LLM API starts returning 50% errors at 3 AM. Your on-call is asleep and users are getting silent blank responses. How do you implement a circuit breaker pattern so the system degrades gracefully instead of silently dying?
What you need to know
The three states
| State | What happens | Moves to |
|---|---|---|
| Closed | Calls go through; results are counted | Open, when the error rate crosses the threshold |
| Open | Calls fail fast without touching the provider | Half-open, after a cool-down (e.g. 60 s) |
| Half-open | A few probe calls go through | Closed if they succeed, Open if they fail |
Trip on error rate over a rolling window — for example more than 50% of at least 20 calls in 30 seconds. Counting consecutive failures is noisy at low 3 AM traffic, where three unlucky calls in a row can open the breaker. The minimum call count stops one failure out of two calls from counting as "50%".
1import time2from collections import deque34class Breaker:5 def __init__(self, window=30, min_calls=20, max_error_rate=0.5, cooldown=60):6 self.window, self.min_calls = window, min_calls7 self.max_error_rate, self.cooldown = max_error_rate, cooldown8 self.calls = deque() # (timestamp, ok)9 self.opened_at = None1011 def allow(self) -> bool:12 if self.opened_at is None:13 return True14 return time.time() - self.opened_at >= self.cooldown # half-open probe1516 def record(self, ok: bool):17 now = time.time()18 if self.opened_at is not None: # this call was a half-open probe19 self.opened_at = None if ok else now # close, or open again20 self.calls.clear()21 return22 self.calls.append((now, ok))23 while self.calls and self.calls[0][0] < now - self.window:24 self.calls.popleft()25 errors = sum(1 for _, good in self.calls if not good)26 if len(self.calls) >= self.min_calls and errors / len(self.calls) > self.max_error_rate:27 self.opened_at = now # tripThis sketch shows the idea. In production, use your gateway's resilience features or a maintained library, and check how it counts failures: some, like pybreaker, count consecutive failures rather than a rate.
What "open" does: the degradation ladder
The key design decision is what the user gets when the breaker is open.
- Secondary provider — a different vendor or region, with its own tested prompt, because prompts do not transfer perfectly.
- Smaller or self-hosted model — lower quality, but an answer.
- Retrieval-only — "I can't write an answer right now; here are the three most relevant documents."
- Honest error — a clear message and a retry time. Never an empty bubble.
The protections around the breaker
- Timeouts on every call. A hung request is worse than an error, because it holds a connection and the user waits.
- Bounded retries with jitter. One or two retries for short glitches, with random delay so clients do not all retry at the same moment. Unbounded retries during an outage multiply load on a sick provider.
- Bulkheads. A separate connection pool per dependency, so one slow provider cannot use up every connection.
Fix the 3 AM part
This is an alerting gap as much as a code gap. Page on the breaker opening and on the share of answers served by fallback. Run a synthetic canary — a fixed test question every minute — so you detect failure even with no users online. Track error rate by provider and, most importantly, user-visible success rate.
A real-life example
Scenario, numbers made up. A food-delivery app's support bot uses one hosted model. At 3:10 AM the provider has a regional incident and 50% of calls fail. The frontend treats an empty body as success, so users see blank replies for two hours. Nobody is paged, because the HTTP status from the app's own API was still 200.
After the fix, the next incident plays out differently. The breaker opens within 30 seconds, traffic moves to a second provider, and the "served by fallback" alert pages on-call, who confirms the provider's status page and goes back to sleep. User-visible success stays at 98% through the incident, and the half-open probes close the breaker automatically when the provider recovers.
Follow-up questions to expect
- "Isn't a retry enough?" — Retries fix short, random failures. During an outage they add load and make users wait longer; the breaker is what stops calling a dependency that is down.
- "How do you choose the thresholds?" — From your normal error rate and traffic. Set the minimum call count so low traffic cannot trip it, and test the settings by injecting failures in staging.
- "What if the stream fails halfway through an answer?" — Send an error event, keep the partial text marked incomplete, and offer to retry on the fallback provider.