Enterprise AI Solutions Architecture

Course Content

Enterprise AI Solutions Architecture

13 sections · 29 lessons

Routing, Fallbacks and Multi-Model Topologies


On a Tuesday during the pilot, Meridian's hosted provider had a partial outage. For about 40 minutes, roughly 30% of calls failed or hung. The pilot had one model and no plan. It retried every failed call three times, which tripled the load on a provider already struggling. Staff on customer calls watched a spinner for 25 seconds before seeing a generic error. Several gave up on the assistant for the rest of the week.

Nothing about that outage was unusual. Every provider has incidents. What made it painful was a design that assumed one model would always be there, and that had no idea what "working less well" should look like.

This lesson designs Meridian's multi-model topology: which model does which job, what happens when one fails, and how much capacity to reserve. It produces part B of MER-05.

Policy answers: the fallback chainProvider A, mid-tier — primary, 88% on golden setProvider B — evaluated at85%, banner shown, up to 4 hoursSearch mode — top threepassages, no generated answerManual PolicyHub search
Every step down is either evaluated like production or is not a model at all; letters skip straight to template mode.

Route by task

Different tasks have different quality bars, so they need different models. The cheapest model that meets a task's bar on your evaluation is the right one; anything more capable is paying for quality you are not using.

TaskModelWhyApproximate cost per call
Classify the input and check for attacksSmall, self-hostedSimple labels, runs on every request, needs under 150 msFixed GPU cost
Rewrite a question into a search querySmall, self-hostedShort output, easy to evaluateFixed GPU cost
Answer a policy questionMid-tier, hostedNeeds careful reading of conditions2.4 cents
Summarise an accountMid-tier, hostedLong input, must not omit3.6 cents
Write letter explanationsMid-tier, hostedTone and clarity6 cents per draft, with its check pass

The team did test a larger, more expensive model for letter wording. It scored two points higher on the tone rubric, which was inside the noise for 100 letters, at about four times the cost. The mid-tier model stayed.

Some systems choose a model per request, sending "easy" questions to a small model and "hard" ones to a large model. This can save money at high volume, but every route is a separate production path that needs its own evaluation, and the classifier that decides "easy" becomes a new way to fail. At Meridian's volume the saving was a few hundred dollars a month, so routing is static, by task. The decision record names the volume at which to revisit it.

Fallbacks that degrade safely

The most important rule in this lesson: a fallback path is a production path. If an untested model takes over during an outage, you have launched an unvalidated system at the worst possible moment. Every fallback is evaluated with the same gates as the primary, or it is not a model at all but a simpler, non-AI mode.

CapabilityFirst choiceFallbackDegraded modeLast resort
Policy answersProvider A, mid-tierProvider B, evaluated at 85% against 88%Search mode: top three passages, no generated answerManual PolicyHub search
Account summaryPrecomputed at case openRegenerate on demandFacts only: tables from records, no proseManual review
Letter draftProvider A, mid-tierNoneTemplate mode: figures filled, explanations blankManual drafting

Letters have no second model on purpose. Provider B was never validated for letter wording, letters are only 70 a day, and none is urgent to the minute. Template mode, with the calculator's figures already filled in, still saves staff most of the work. For policy answers, provider B is acceptable for up to four hours, with a banner in the panel saying a backup model is in use; after that, the service moves to search mode until provider A returns.

A router with a circuit breaker

A circuit breaker stops calling a dependency that keeps failing, and tries it again after a pause. It protects users from waiting on a dead route and protects the struggling provider from retry storms. The code below is the core of Meridian's router, behind a small LLM interface that any provider adapter implements.

Python
import timefrom dataclasses import dataclass, fieldfrom typing import Callable, Protocolclass LLM(Protocol):    def complete(self, model: str, messages: list[dict], max_tokens: int, timeout: float) -> str: ...@dataclassclass Breaker:    threshold: int = 5        # consecutive failures before opening    cool_off: float = 30.0    # seconds before trying again    failures: int = 0    opened_at: float | None = None    def allow(self) -> bool:        if self.opened_at is None:            return True        if time.monotonic() - self.opened_at < self.cool_off:            return False        self.opened_at, self.failures = None, self.threshold - 1  # half-open: one trial call        return True    def record(self, ok: bool) -> None:        self.failures = 0 if ok else self.failures + 1        if self.failures >= self.threshold:            self.opened_at = time.monotonic()@dataclassclass Route:    name: str    client: LLM    model: str    timeout: float    breaker: Breaker = field(default_factory=Breaker)def complete(routes: list[Route], messages: list[dict], degrade: Callable[[], str]) -> tuple[str, str]:    for route in routes:        if not route.breaker.allow():            continue        try:            text = route.client.complete(route.model, messages, max_tokens=600, timeout=route.timeout)            route.breaker.record(True)            return route.name, text        except Exception:            route.breaker.record(False)    return "degraded", degrade()

complete tries each route in order and skips any whose breaker is open. After five consecutive failures a breaker opens for 30 seconds, then lets one trial call through; if that call fails, it opens again at once. If every route is closed or failing, degrade returns the degraded mode, such as search results. The function returns the route name, so the panel can show a banner and the trace can record which path answered.

Notice what is missing: retries on the same route. On an interactive path, a failed call has usually already cost the user seconds. Retrying it in place adds more seconds and more load. The breaker means only the first few users during an outage pay the price of a timeout; after that, traffic goes straight to the fallback.

Capacity and rate limits

Providers limit requests and tokens per minute. Size those limits for peak, not average. Meridian's peak is between 11:00 and 13:00, at about three times the average rate: roughly 11 policy answers a minute, or about 60,000 input tokens a minute. The bank asked for a limit of 200,000 input tokens a minute, over three times peak, for headroom.

Two background loads are easy to forget. First, precompute bursts: on Monday mornings, up to 150 hardship cases created from weekend web forms open within the first hour, each triggering a summary. These go through a queue that drains at a fixed rate, so they cannot starve interactive users. Second, evaluation runs: running 300 golden questions against three configurations uses real quota. Evaluation runs use a separate project with its own limits, so a release test at 11:30 never slows a staff member on a call.

The routing table, fallback chains and capacity figures above are part B of MER-05.

Check your understanding

0 of 3 answered

1.During an outage, why is switching to an untested model a poor fallback?

2.Why does Meridian's router not retry a failed call on the same route?

3.Evaluation runs of 300 questions use a separate project with its own rate limits. What problem does that prevent?