Scenario-Based AI Engineering Questions

Course Content

Scenario-Based AI Engineering Questions

26 sections · 146 lessons

Your entire product depends on one LLM provider. They have a 4-hour outage during your launch, then raise prices 3×. How do you build LLM-provider-agnostic infrastructure without losing quality?


One call path with a proven fallbackInternalgenerate()adapterPer-providerprompt variantCircuit breakeron the primarySecondary:2 percentlive trafficDegraded answer,never an error pageA 400-case golden eval decides when a provider is ready.
A fallback that carries real traffic every day is the only kind that works during an outage.

What you need to know

One boundary in the code

All model calls go through one internal function, for example generate(messages, tools, schema, budget), built on a gateway library such as LiteLLM or a thin adapter of your own. No provider SDK is imported anywhere else. Keep provider-specific wins available as optional flags — prompt caching, structured output, extended thinking — with fallbacks when a provider lacks them. Abstracting down to the lowest common feature set makes every provider equally mediocre.

Prompts do not port

The same instructions behave differently across model families: formatting, tool-calling style and refusal thresholds differ. Keep one logical prompt name with a variant per provider, and let the eval suite decide when a variant is ready.

Failover that actually works

Python
PROVIDERS = [("primary", call_primary), ("secondary", call_secondary)]async def generate(req):    order = PROVIDERS if random.random() > 0.02 else PROVIDERS[::-1]   # 2% live on secondary    for name, call in order:        if breaker[name].is_open():            continue        try:            reply = await asyncio.wait_for(call(req, prompt=PROMPTS[req.task][name]), timeout=20)            breaker[name].success()            return reply        except (TimeoutError, ProviderServerError):            breaker[name].failure()    return degraded_answer(req)            # cached answer or clear message, never a stack trace

The permanent 2% on the secondary is the important line. A fallback path that never runs is usually broken when you need it.

OptionProtects againstCost
Second commercial providerOutages, price risesPrompt variants, eval runs
Same model via a second cloudProvider-region outageFewer model choices
Self-hosted open modelPrice rises, residencyGPUs and operations work

A real-life example

Scenario, numbers made up. A legal-drafting startup launches on a Monday; its only provider is down for four hours. A month later, the provider's price for the model they use triples.

The team adds an adapter, a 400-case golden eval, and a second provider with its own prompt variants. The second provider scores 3 points lower on the eval out of the box and within 1 point after two prompt revisions. They keep 2% of traffic on it. For clause classification, 60% of calls, a self-hosted open model reaches parity on the eval. At the next outage, the breaker opens in 30 seconds and users see slightly slower answers instead of errors, and the price negotiation goes differently when the team can show a working alternative.

Follow-up questions to expect

  • "Won't the abstraction slow you down?" — Only if it hides features. Expose capabilities as flags, and let teams use provider-specific features with a fallback defined.
  • "How do you keep outputs consistent across providers?" — Enforce structure with schemas and validators, and judge equivalence on the golden set, not by reading a few outputs.
  • "Do you fail over mid-conversation?" — Yes, if the conversation history is stored in your own format; that is another reason not to rely on provider-side conversation state.