Course Content
Enterprise AI Solutions Architecture
13 sections · 29 lessons
Routing, Fallbacks and Multi-Model Topologies
On a Tuesday during the pilot, Meridian's hosted provider had a partial outage. For about 40 minutes, roughly 30% of calls failed or hung. The pilot had one model and no plan. It retried every failed call three times, which tripled the load on a provider already struggling. Staff on customer calls watched a spinner for 25 seconds before seeing a generic error. Several gave up on the assistant for the rest of the week.
Nothing about that outage was unusual. Every provider has incidents. What made it painful was a design that assumed one model would always be there, and that had no idea what "working less well" should look like.
This lesson designs Meridian's multi-model topology: which model does which job, what happens when one fails, and how much capacity to reserve. It produces part B of MER-05.
Route by task
Different tasks have different quality bars, so they need different models. The cheapest model that meets a task's bar on your evaluation is the right one; anything more capable is paying for quality you are not using.
| Task | Model | Why | Approximate cost per call |
|---|---|---|---|
| Classify the input and check for attacks | Small, self-hosted | Simple labels, runs on every request, needs under 150 ms | Fixed GPU cost |
| Rewrite a question into a search query | Small, self-hosted | Short output, easy to evaluate | Fixed GPU cost |
| Answer a policy question | Mid-tier, hosted | Needs careful reading of conditions | 2.4 cents |
| Summarise an account | Mid-tier, hosted | Long input, must not omit | 3.6 cents |
| Write letter explanations | Mid-tier, hosted | Tone and clarity | 6 cents per draft, with its check pass |
The team did test a larger, more expensive model for letter wording. It scored two points higher on the tone rubric, which was inside the noise for 100 letters, at about four times the cost. The mid-tier model stayed.
Some systems choose a model per request, sending "easy" questions to a small model and "hard" ones to a large model. This can save money at high volume, but every route is a separate production path that needs its own evaluation, and the classifier that decides "easy" becomes a new way to fail. At Meridian's volume the saving was a few hundred dollars a month, so routing is static, by task. The decision record names the volume at which to revisit it.
Fallbacks that degrade safely
The most important rule in this lesson: a fallback path is a production path. If an untested model takes over during an outage, you have launched an unvalidated system at the worst possible moment. Every fallback is evaluated with the same gates as the primary, or it is not a model at all but a simpler, non-AI mode.
| Capability | First choice | Fallback | Degraded mode | Last resort |
|---|---|---|---|---|
| Policy answers | Provider A, mid-tier | Provider B, evaluated at 85% against 88% | Search mode: top three passages, no generated answer | Manual PolicyHub search |
| Account summary | Precomputed at case open | Regenerate on demand | Facts only: tables from records, no prose | Manual review |
| Letter draft | Provider A, mid-tier | None | Template mode: figures filled, explanations blank | Manual drafting |
Letters have no second model on purpose. Provider B was never validated for letter wording, letters are only 70 a day, and none is urgent to the minute. Template mode, with the calculator's figures already filled in, still saves staff most of the work. For policy answers, provider B is acceptable for up to four hours, with a banner in the panel saying a backup model is in use; after that, the service moves to search mode until provider A returns.
A router with a circuit breaker
A circuit breaker stops calling a dependency that keeps failing, and tries it again after a pause. It protects users from waiting on a dead route and protects the struggling provider from retry storms. The code below is the core of Meridian's router, behind a small LLM interface that any provider adapter implements.
1import time2from dataclasses import dataclass, field3from typing import Callable, Protocol45class LLM(Protocol):6 def complete(self, model: str, messages: list[dict], max_tokens: int, timeout: float) -> str: ...78@dataclass9class Breaker:10 threshold: int = 5 # consecutive failures before opening11 cool_off: float = 30.0 # seconds before trying again12 failures: int = 013 opened_at: float | None = None1415 def allow(self) -> bool:16 if self.opened_at is None:17 return True18 if time.monotonic() - self.opened_at < self.cool_off:19 return False20 self.opened_at, self.failures = None, self.threshold - 1 # half-open: one trial call21 return True2223 def record(self, ok: bool) -> None:24 self.failures = 0 if ok else self.failures + 125 if self.failures >= self.threshold:26 self.opened_at = time.monotonic()2728@dataclass29class Route:30 name: str31 client: LLM32 model: str33 timeout: float34 breaker: Breaker = field(default_factory=Breaker)3536def complete(routes: list[Route], messages: list[dict], degrade: Callable[[], str]) -> tuple[str, str]:37 for route in routes:38 if not route.breaker.allow():39 continue40 try:41 text = route.client.complete(route.model, messages, max_tokens=600, timeout=route.timeout)42 route.breaker.record(True)43 return route.name, text44 except Exception:45 route.breaker.record(False)46 return "degraded", degrade()complete tries each route in order and skips any whose breaker is open. After five consecutive failures a breaker opens for 30 seconds, then lets one trial call through; if that call fails, it opens again at once. If every route is closed or failing, degrade returns the degraded mode, such as search results. The function returns the route name, so the panel can show a banner and the trace can record which path answered.
Notice what is missing: retries on the same route. On an interactive path, a failed call has usually already cost the user seconds. Retrying it in place adds more seconds and more load. The breaker means only the first few users during an outage pay the price of a timeout; after that, traffic goes straight to the fallback.
Capacity and rate limits
Providers limit requests and tokens per minute. Size those limits for peak, not average. Meridian's peak is between 11:00 and 13:00, at about three times the average rate: roughly 11 policy answers a minute, or about 60,000 input tokens a minute. The bank asked for a limit of 200,000 input tokens a minute, over three times peak, for headroom.
Two background loads are easy to forget. First, precompute bursts: on Monday mornings, up to 150 hardship cases created from weekend web forms open within the first hour, each triggering a summary. These go through a queue that drains at a fixed rate, so they cannot starve interactive users. Second, evaluation runs: running 300 golden questions against three configurations uses real quota. Evaluation runs use a separate project with its own limits, so a release test at 11:30 never slows a staff member on a call.
The routing table, fallback chains and capacity figures above are part B of MER-05.
Check your understanding
0 of 3 answered
1.During an outage, why is switching to an untested model a poor fallback?
2.Why does Meridian's router not retry a failed call on the same route?
3.Evaluation runs of 300 questions use a separate project with its own rate limits. What problem does that prevent?