Course Content
Scenario-Based AI Engineering Questions
26 sections · 146 lessons
Your entire product depends on one LLM provider. They have a 4-hour outage during your launch, then raise prices 3×. How do you build LLM-provider-agnostic infrastructure without losing quality?
What you need to know
One boundary in the code
All model calls go through one internal function, for example generate(messages, tools, schema, budget), built on a gateway library such as LiteLLM or a thin adapter of your own. No provider SDK is imported anywhere else. Keep provider-specific wins available as optional flags — prompt caching, structured output, extended thinking — with fallbacks when a provider lacks them. Abstracting down to the lowest common feature set makes every provider equally mediocre.
Prompts do not port
The same instructions behave differently across model families: formatting, tool-calling style and refusal thresholds differ. Keep one logical prompt name with a variant per provider, and let the eval suite decide when a variant is ready.
Failover that actually works
1PROVIDERS = [("primary", call_primary), ("secondary", call_secondary)]23async def generate(req):4 order = PROVIDERS if random.random() > 0.02 else PROVIDERS[::-1] # 2% live on secondary5 for name, call in order:6 if breaker[name].is_open():7 continue8 try:9 reply = await asyncio.wait_for(call(req, prompt=PROMPTS[req.task][name]), timeout=20)10 breaker[name].success()11 return reply12 except (TimeoutError, ProviderServerError):13 breaker[name].failure()14 return degraded_answer(req) # cached answer or clear message, never a stack traceThe permanent 2% on the secondary is the important line. A fallback path that never runs is usually broken when you need it.
| Option | Protects against | Cost |
|---|---|---|
| Second commercial provider | Outages, price rises | Prompt variants, eval runs |
| Same model via a second cloud | Provider-region outage | Fewer model choices |
| Self-hosted open model | Price rises, residency | GPUs and operations work |
A real-life example
Scenario, numbers made up. A legal-drafting startup launches on a Monday; its only provider is down for four hours. A month later, the provider's price for the model they use triples.
The team adds an adapter, a 400-case golden eval, and a second provider with its own prompt variants. The second provider scores 3 points lower on the eval out of the box and within 1 point after two prompt revisions. They keep 2% of traffic on it. For clause classification, 60% of calls, a self-hosted open model reaches parity on the eval. At the next outage, the breaker opens in 30 seconds and users see slightly slower answers instead of errors, and the price negotiation goes differently when the team can show a working alternative.
Follow-up questions to expect
- "Won't the abstraction slow you down?" — Only if it hides features. Expose capabilities as flags, and let teams use provider-specific features with a fallback defined.
- "How do you keep outputs consistent across providers?" — Enforce structure with schemas and validators, and judge equivalence on the golden set, not by reading a few outputs.
- "Do you fail over mid-conversation?" — Yes, if the conversation history is stored in your own format; that is another reason not to rely on provider-side conversation state.