Enterprise AI Solutions Architecture

Course Content

Enterprise AI Solutions Architecture

13 sections · 29 lessons

Trust Your Own Tests: Benchmarks, Demos and Vendor Promises


Three weeks into discovery, a vendor presented to Meridian's procurement team. Their banking assistant was "98% accurate, hallucination-free, and top of the leaderboard on a leading reasoning benchmark." Procurement liked the numbers. The Head of Servicing liked the demo. The architect asked one question: "98% of what, graded by whom?"

That question is the core of this lesson. Every AI decision rests on claims: about models, about products, about your own data and your own users. Some claims are backed by strong evidence. Many are not. The architect's job is not to be cynical; it is to know which is which, and to write it down so that nobody mistakes an assumption for a fact.

The result is the second part of MER-02: an evidence register listing every claim the design depends on, how strong the evidence is today, and how strong it must be before each decision.

The evidence ladder, E0 at the bottomE0 claim: slidesand sales callsE1 publicbenchmarkE2 externalstudy, similar taskE3 our data,our gradersE4 measuredwith our own userstopbottomThe vendor's 98% became 81% once Meridian's policy team did the grading.
Match the evidence to the decision: shortlist on E1, commit the architecture on E3, scale to 450 staff on E4.

What public benchmarks measure

A public benchmark is a fixed set of tasks with known answers, used to compare models. There are benchmarks for general knowledge, maths, coding, long-document reading, tool use and many more. They are useful for one thing: shortlisting. If a model does badly on long-document reading, it is unlikely to do well on your 2,100 pages of policy.

They are weak evidence for anything more, for five reasons.

  • Contamination. Test questions leak onto the internet and into training data. A model may have seen the answers.
  • Distribution mismatch. The benchmark is not your documents, your questions or your users. Meridian's staff ask about product codes and internal procedures that no public dataset contains.
  • Metric mismatch. Many benchmarks score multiple choice. Your task is a cited, complete answer that keeps every condition. Those are different skills.
  • Saturation. Top models often score within a few points of each other. The differences are smaller than the differences your own prompt and retrieval design will make.
  • Tuning to the test. Once a benchmark matters commercially, vendors optimise for it. The score rises faster than real-world ability.

A leaderboard position tells you a model is capable in general. It tells you almost nothing about whether it will answer Meridian's arrears questions correctly.

Reading a vendor claim

A number like "98% accurate" is not wrong or right on its own. It is incomplete. Five questions make it complete.

QuestionWhy it mattersA good answer sounds like
Accurate at which task?Accuracy on FAQs is not accuracy on edge cases"Cited answers to servicing policy questions"
On whose data?Their documents are cleaner than yours"On 300 questions from two client banks, method attached"
Graded how, by whom?Lenient graders inflate scores"Two domain experts, partial answers count as wrong"
How many cases, what spread?98% on 50 cases is a wide range"400 cases, interval 96% to 99%"
Which model version, and will we get it?The tested version may not be the one you buy"Pinned version, and we notify you before changing it"

Treat some phrases as warning signs. "Hallucination-free" cannot be true of any current generative system; retrieval and checks reduce invented content, but nobody can promise zero. "Enterprise-grade accuracy" has no definition. "Validated by banks" means nothing without the method.

The evidence ladder

Not all evidence is equal. Meridian uses a five-level scale, and every claim in the design is graded on it.

LevelNameExample
E0ClaimA vendor slide, a sales call, a blog post
E1Public benchmarkLeaderboard scores on general tasks
E2External evaluation on a similar taskA published study with its method, a peer bank's shared results
E3Our offline evaluationOur questions, our documents, our graders
E4Measured with our usersA shadow run or pilot in production conditions

The rule is to match the evidence to the size of the decision. Shortlisting three models for a spike can rest on E1. Committing the architecture needs E3. Going live needs E3 plus a pilot at E4. Scaling to all 450 staff needs E4 measured over weeks, not days. Big, hard-to-reverse decisions need high rungs; small, cheap ones can use low rungs.

Getting to E3 is cheaper than most teams expect, if the comparison is fair. Meridian's rules for any bake-off are short. Every candidate gets the same questions, the same retrieved passages and the same output format, so the only thing that varies is the thing being compared. Graders see answers without knowing which candidate wrote them. Every candidate is run at least twice, so the team sees how much each one varies. And the grading guide is written before the first run, so nobody adjusts the standard after seeing which candidate they prefer. A fair bake-off on 150 questions costs a few days of policy-team time and less than a hundred dollars in model calls.

The evidence register

The register lists every claim the design depends on. "Assumptions are allowed; hidden assumptions are not." A reviewer who sees an E0 assumption, labelled as such, with a plan to raise it, will usually accept it. A reviewer who discovers it on their own will doubt everything else.

IDClaim the design depends onSourceNowNeededHow we raise itOwner
EV-01A mid-tier hosted model with retrieval answers policy questions at 85% or betterVendor and public benchmarksE1E3 before design commit150-question spikeArchitect
EV-02PolicyHub documents parse with tables intactAssumptionE0E3Parse the 30 hardest documentsEng lead
EV-03Ledger API serves 5 requests a second at p95 under 300 msSystem owner's 2024 load testE2E3Load test in the test environmentIntegration lead
EV-04Provider keeps no prompts and does not train on themDraft contractE0Contract signedLegal and procurement reviewProcurement
EV-05Staff trust cited answers enough to use themHead of Servicing's viewE0E430-person pilot, usage and ratingsHead of Servicing
EV-06Better letter explanations reduce complaintsHypothesisE0E4Compare complaint rates over six monthsOps QA lead

Look at the "Now" and "Needed" columns together: the gap between them is the discovery plan. Section 3 turns these gaps into a timed spike with clear stop criteria.

Using the tables on one decision

The three tables are easier to use once you have watched one real decision pass through them. Take the biggest decision of Meridian's discovery: should the design commit to a mid-tier hosted model with retrieval for policy answers? That is claim EV-01 in the register.

Start with the evidence ladder and ask what the bank actually knew that week. A leaderboard showed the shortlisted models reading long documents well: E1. The vendor's "98%" and the demo that impressed the Head of Servicing were E0. A demo proves that something can work once, on questions someone chose in advance, and nothing more. Then ask how big the decision is. Committing the architecture shapes months of build and is hard to reverse, so the rule says E3. The distance from E1 to E3 is the work still to do.

Next, turn the five vendor questions around and use them as the specification for your own test. Which task: cited answers to servicing policy questions. Whose data: Meridian's documents and 150 of its own questions, 30 of them written to be hard. Graded how: two policy analysts who cannot see which candidate wrote each answer, with a missing condition counted as wrong. How many cases: 150, which gives a range about eleven points wide. Which version: a pinned model version, recorded next to the result. Every answer you would demand from a vendor is a design choice in your own bake-off.

Finally, write the result back into the register. The spike in section 3 scored 83%, with a range of 77% to 88%, so EV-01 moved from E1 to E3 and the architecture was committed. On the same day EV-05, staff trust, was still at E0, and that was acceptable: it was labelled, it had an owner, and the pilot was its plan. A reviewer reading both rows can see exactly which parts of the design rest on evidence and which still rest on belief.

Check your understanding

0 of 3 answered

1.A model tops a public reasoning leaderboard. What is the strongest conclusion Meridian can draw for its policy assistant?

2.The vendor's 98% fell to 81% on Meridian's questions. Which difference mattered most?

3.Which decision needs E4 evidence, measured with Meridian's own users?