Scenario-Based AI Engineering Questions

Course Content

Scenario-Based AI Engineering Questions

26 sections · 146 lessons

Your LLM API starts returning 50% errors at 3 AM. Your on-call is asleep and users are getting silent blank responses. How do you implement a circuit breaker pattern so the system degrades gracefully instead of silently dying?


What an open breaker falls through toPrimary provider — open, fail fastSecondary provider, own tested promptSmaller or self-hosted modelRetrieval only — top three documentsHonest error and retry time, never blank
A breaker only helps if opening it leads somewhere; each rung trades quality for the certainty of saying something true.

What you need to know

The three states

StateWhat happensMoves to
ClosedCalls go through; results are countedOpen, when the error rate crosses the threshold
OpenCalls fail fast without touching the providerHalf-open, after a cool-down (e.g. 60 s)
Half-openA few probe calls go throughClosed if they succeed, Open if they fail

Trip on error rate over a rolling window — for example more than 50% of at least 20 calls in 30 seconds. Counting consecutive failures is noisy at low 3 AM traffic, where three unlucky calls in a row can open the breaker. The minimum call count stops one failure out of two calls from counting as "50%".

Python
import timefrom collections import dequeclass Breaker:    def __init__(self, window=30, min_calls=20, max_error_rate=0.5, cooldown=60):        self.window, self.min_calls = window, min_calls        self.max_error_rate, self.cooldown = max_error_rate, cooldown        self.calls = deque()          # (timestamp, ok)        self.opened_at = None    def allow(self) -> bool:        if self.opened_at is None:            return True        return time.time() - self.opened_at >= self.cooldown   # half-open probe    def record(self, ok: bool):        now = time.time()        if self.opened_at is not None:              # this call was a half-open probe            self.opened_at = None if ok else now    # close, or open again            self.calls.clear()            return        self.calls.append((now, ok))        while self.calls and self.calls[0][0] < now - self.window:            self.calls.popleft()        errors = sum(1 for _, good in self.calls if not good)        if len(self.calls) >= self.min_calls and errors / len(self.calls) > self.max_error_rate:            self.opened_at = now                    # trip

This sketch shows the idea. In production, use your gateway's resilience features or a maintained library, and check how it counts failures: some, like pybreaker, count consecutive failures rather than a rate.

What "open" does: the degradation ladder

The key design decision is what the user gets when the breaker is open.

  1. Secondary provider — a different vendor or region, with its own tested prompt, because prompts do not transfer perfectly.
  2. Smaller or self-hosted model — lower quality, but an answer.
  3. Retrieval-only — "I can't write an answer right now; here are the three most relevant documents."
  4. Honest error — a clear message and a retry time. Never an empty bubble.

The protections around the breaker

  • Timeouts on every call. A hung request is worse than an error, because it holds a connection and the user waits.
  • Bounded retries with jitter. One or two retries for short glitches, with random delay so clients do not all retry at the same moment. Unbounded retries during an outage multiply load on a sick provider.
  • Bulkheads. A separate connection pool per dependency, so one slow provider cannot use up every connection.

Fix the 3 AM part

This is an alerting gap as much as a code gap. Page on the breaker opening and on the share of answers served by fallback. Run a synthetic canary — a fixed test question every minute — so you detect failure even with no users online. Track error rate by provider and, most importantly, user-visible success rate.

A real-life example

Scenario, numbers made up. A food-delivery app's support bot uses one hosted model. At 3:10 AM the provider has a regional incident and 50% of calls fail. The frontend treats an empty body as success, so users see blank replies for two hours. Nobody is paged, because the HTTP status from the app's own API was still 200.

After the fix, the next incident plays out differently. The breaker opens within 30 seconds, traffic moves to a second provider, and the "served by fallback" alert pages on-call, who confirms the provider's status page and goes back to sleep. User-visible success stays at 98% through the incident, and the half-open probes close the breaker automatically when the provider recovers.

Follow-up questions to expect

  • "Isn't a retry enough?" — Retries fix short, random failures. During an outage they add load and make users wait longer; the breaker is what stops calling a dependency that is down.
  • "How do you choose the thresholds?" — From your normal error rate and traffic. Set the minimum call count so low traffic cannot trip it, and test the settings by injecting failures in staging.
  • "What if the stream fails halfway through an answer?" — Send an error event, keep the partial text marked incomplete, and offer to retry on the fallback provider.