Course Content
Enterprise AI Solutions Architecture
13 sections · 29 lessons
When the Same Question Gets Different Answers
During Meridian's first pilot, a team lead asked the assistant whether a customer on a payment holiday could also have their interest frozen. She asked the same question three times during a demo, and got the same correct answer three times. The next morning a staff member asked it again, word for word, and got an answer that said yes without any condition. The policy says yes only for customers with a vulnerability flag. Nothing in the code had changed.
Engineers who have built normal services find this uncomfortable, and they should. A normal service is a function: the same input gives the same output, and a failing test fails every time. A model-backed service is closer to a noisy measurement. It has a rate of being right, not a guarantee.
This lesson explains what that means for architecture. The short version: you cannot inspect quality into such a system at the end. You design how it will be measured from the first day, and that measurement is engineering work with its own components, data and gates.
Three properties of a probabilistic service
Variance. The same request can produce different outputs. Setting temperature to zero reduces this but does not remove it. Providers batch requests together on shared hardware, and small numerical differences can change which token wins. Providers also update serving infrastructure. Treat "deterministic at temperature zero" as a hope, not a design property.
Fluent failure. A normal service fails loudly: a status code, an exception, a timeout. A model fails by returning a confident, well-written, wrong answer. Nothing in the response tells you it is wrong. Monitoring built for loud failures, such as error rates and HTTP 500 counts, will show a healthy system while it misleads staff.
Sensitivity to small changes. Rewording one line of the prompt, changing the order of retrieved passages, or moving to a new model version can shift quality by several points, up or down, on cases that seem unrelated to the change. A change that fixes ten failures may create six new ones.
Pass rates need sample sizes
Because quality is a rate, it must be estimated from a sample, and a small sample gives a wide range. If the assistant gets 18 of 20 test questions right, the true rate could reasonably be anywhere from about 70% to 97%. That range is too wide to decide anything.
The Wilson score interval is a standard way to put a range around a pass rate. It behaves better than the textbook formula for small samples and for rates near 100%, which is where AI quality targets usually sit.
1from math import sqrt23def wilson_interval(passed: int, total: int, z: float = 1.96) -> tuple[float, float]:4 """95% confidence interval for a pass rate (Wilson score)."""5 if total == 0:6 raise ValueError("no test cases")7 p = passed / total8 denom = 1 + z * z / total9 centre = (p + z * z / (2 * total)) / denom10 half = z * sqrt(p * (1 - p) / total + z * z / (4 * total * total)) / denom11 return max(0.0, centre - half), min(1.0, centre + half)1213for n in (20, 50, 150, 400):14 lo, hi = wilson_interval(round(0.9 * n), n)15 print(f"{n:>4} cases, 90% observed -> {lo:.0%} to {hi:.0%}")The function takes the number of passing cases and the total, and returns the low and high end of the range that probably contains the true rate. Running the loop prints the table below. Each row shows the same observed 90% on a larger test set.
| Test cases | Observed | Probable true rate |
|---|---|---|
| 20 | 90% | 70% to 97% |
| 50 | 90% | 79% to 96% |
| 150 | 90% | 84% to 94% |
| 400 | 90% | 87% to 93% |
Two practical rules follow. First, write thresholds as a lower bound, not a point: "the 95% interval's low end is at least 85%" is a requirement a reviewer can check. Second, a change that moves the score by two points on 150 cases is noise, not progress. To compare two prompt versions, run both on the same cases and count the cases where they disagree; that paired view is far more sensitive than comparing two totals.
Evaluation is engineering
Once quality is a rate, the evaluation set becomes the most important asset in the project. It is three things at once.
- The executable specification. Each case says, "for this input, a good output has these properties." Written with the policy team, it is the most precise statement of what "correct" means that the bank has ever had.
- The test suite. It runs on every prompt, model or index change, just as unit tests run on every commit.
- The release gate. A change that lowers the lower bound below the threshold does not ship.
Treat it like code. It lives in version control. Changes to it are reviewed. It has an owner, coverage targets per product and per policy area, and a changelog. Section 8 designs the full evaluation architecture; for now the point is that it must be planned and budgeted from the first week, not written in the last.
Sample it or enforce it
Not every requirement should be a rate. Some failures are so costly that "98% of the time" is unacceptable. For those, the design must move the work out of the model and into something deterministic that code can check every time.
Measure by sampling
- Is the policy answer correct and complete?
- Does the summary mention every earlier hardship plan?
- Is the letter's tone clear and kind?
- Would a senior agent accept the draft?
Enforce with code
- Do letter figures equal the calculator output?
- Are the mandatory paragraphs present, word for word?
- Does every citation point to a real, current policy section?
- Is the customer in the draft the customer on the case?
The left column needs judgement, so you sample it and put a range around the result. The right column has a single correct answer that code can check on every output before a human sees it. A good design moves as many requirements as possible from left to right. This is one of the architect's highest-value moves, and it is why Meridian's letters take figures from the Hardship Calculator.
| ID | Hypothesis | Threshold | Method | Owner |
|---|---|---|---|---|
| Q-01 | Policy answers are correct and cited | 95% interval low end at least 85% | 150 questions graded by the policy team | Policy lead |
| Q-02 | Summaries include every material event | 98% on 100 hardship cases | Checklist grading by senior agents | Ops QA lead |
| Q-03 | Letter figures match the calculator | 100%, every draft | Code check before display | Eng lead |
| Q-04 | Mandatory paragraphs are present and unaltered | 100%, every draft | Code check before display | Eng lead |
| Q-05 | Drafts accepted with minor or no edits | At least 70% | Edit distance and staff rating in pilot | Head of Servicing |
Q-03 and Q-04 are not statistical at all. They are guarantees, and the design will have to earn them.
Check your understanding
0 of 3 answered
1.The assistant passes 18 of 20 policy questions. A manager says, "90% accurate, ship it." What is the best response?
2.Which Meridian requirement should be enforced by code on every output rather than measured by sampling?
3.Version B of a prompt scores 91% and version A scores 89%, each on the same 150 cases. What should you conclude?