Enterprise AI Solutions Architecture

Course Content

Enterprise AI Solutions Architecture

13 sections · 29 lessons

When the Same Question Gets Different Answers


During Meridian's first pilot, a team lead asked the assistant whether a customer on a payment holiday could also have their interest frozen. She asked the same question three times during a demo, and got the same correct answer three times. The next morning a staff member asked it again, word for word, and got an answer that said yes without any condition. The policy says yes only for customers with a vulnerability flag. Nothing in the code had changed.

Engineers who have built normal services find this uncomfortable, and they should. A normal service is a function: the same input gives the same output, and a failing test fails every time. A model-backed service is closer to a noisy measurement. It has a rate of being right, not a guarantee.

This lesson explains what that means for architecture. The short version: you cannot inspect quality into such a system at the end. You design how it will be measured from the first day, and that measurement is engineering work with its own components, data and gates.

The same 90% score on four test-set sizes90%70% to 97%90%79% to 96%90%84% to 94%90%87% to 93%ObservedProbable true rate20 cases50 cases150 cases400 casesWilson score interval at 95% confidence.
Eighteen out of twenty proves almost nothing; write thresholds as the low end of a range, not as a single score.

Three properties of a probabilistic service

Variance. The same request can produce different outputs. Setting temperature to zero reduces this but does not remove it. Providers batch requests together on shared hardware, and small numerical differences can change which token wins. Providers also update serving infrastructure. Treat "deterministic at temperature zero" as a hope, not a design property.

Fluent failure. A normal service fails loudly: a status code, an exception, a timeout. A model fails by returning a confident, well-written, wrong answer. Nothing in the response tells you it is wrong. Monitoring built for loud failures, such as error rates and HTTP 500 counts, will show a healthy system while it misleads staff.

Sensitivity to small changes. Rewording one line of the prompt, changing the order of retrieved passages, or moving to a new model version can shift quality by several points, up or down, on cases that seem unrelated to the change. A change that fixes ten failures may create six new ones.

Pass rates need sample sizes

Because quality is a rate, it must be estimated from a sample, and a small sample gives a wide range. If the assistant gets 18 of 20 test questions right, the true rate could reasonably be anywhere from about 70% to 97%. That range is too wide to decide anything.

The Wilson score interval is a standard way to put a range around a pass rate. It behaves better than the textbook formula for small samples and for rates near 100%, which is where AI quality targets usually sit.

Python
from math import sqrtdef wilson_interval(passed: int, total: int, z: float = 1.96) -> tuple[float, float]:    """95% confidence interval for a pass rate (Wilson score)."""    if total == 0:        raise ValueError("no test cases")    p = passed / total    denom = 1 + z * z / total    centre = (p + z * z / (2 * total)) / denom    half = z * sqrt(p * (1 - p) / total + z * z / (4 * total * total)) / denom    return max(0.0, centre - half), min(1.0, centre + half)for n in (20, 50, 150, 400):    lo, hi = wilson_interval(round(0.9 * n), n)    print(f"{n:>4} cases, 90% observed -> {lo:.0%} to {hi:.0%}")

The function takes the number of passing cases and the total, and returns the low and high end of the range that probably contains the true rate. Running the loop prints the table below. Each row shows the same observed 90% on a larger test set.

Test casesObservedProbable true rate
2090%70% to 97%
5090%79% to 96%
15090%84% to 94%
40090%87% to 93%

Two practical rules follow. First, write thresholds as a lower bound, not a point: "the 95% interval's low end is at least 85%" is a requirement a reviewer can check. Second, a change that moves the score by two points on 150 cases is noise, not progress. To compare two prompt versions, run both on the same cases and count the cases where they disagree; that paired view is far more sensitive than comparing two totals.

Evaluation is engineering

Once quality is a rate, the evaluation set becomes the most important asset in the project. It is three things at once.

  • The executable specification. Each case says, "for this input, a good output has these properties." Written with the policy team, it is the most precise statement of what "correct" means that the bank has ever had.
  • The test suite. It runs on every prompt, model or index change, just as unit tests run on every commit.
  • The release gate. A change that lowers the lower bound below the threshold does not ship.

Treat it like code. It lives in version control. Changes to it are reviewed. It has an owner, coverage targets per product and per policy area, and a changelog. Section 8 designs the full evaluation architecture; for now the point is that it must be planned and budgeted from the first week, not written in the last.

Sample it or enforce it

Not every requirement should be a rate. Some failures are so costly that "98% of the time" is unacceptable. For those, the design must move the work out of the model and into something deterministic that code can check every time.

Measure by sampling

  • Is the policy answer correct and complete?
  • Does the summary mention every earlier hardship plan?
  • Is the letter's tone clear and kind?
  • Would a senior agent accept the draft?

Enforce with code

  • Do letter figures equal the calculator output?
  • Are the mandatory paragraphs present, word for word?
  • Does every citation point to a real, current policy section?
  • Is the customer in the draft the customer on the case?

The left column needs judgement, so you sample it and put a range around the result. The right column has a single correct answer that code can check on every output before a human sees it. A good design moves as many requirements as possible from left to right. This is one of the architect's highest-value moves, and it is why Meridian's letters take figures from the Hardship Calculator.

IDHypothesisThresholdMethodOwner
Q-01Policy answers are correct and cited95% interval low end at least 85%150 questions graded by the policy teamPolicy lead
Q-02Summaries include every material event98% on 100 hardship casesChecklist grading by senior agentsOps QA lead
Q-03Letter figures match the calculator100%, every draftCode check before displayEng lead
Q-04Mandatory paragraphs are present and unaltered100%, every draftCode check before displayEng lead
Q-05Drafts accepted with minor or no editsAt least 70%Edit distance and staff rating in pilotHead of Servicing

Q-03 and Q-04 are not statistical at all. They are guarantees, and the design will have to earn them.

Check your understanding

0 of 3 answered

1.The assistant passes 18 of 20 policy questions. A manager says, "90% accurate, ship it." What is the best response?

2.Which Meridian requirement should be enforced by code on every output rather than measured by sampling?

3.Version B of a prompt scores 91% and version A scores 89%, each on the same 150 cases. What should you conclude?