Course Content
Enterprise AI Solutions Architecture
13 sections · 29 lessons
The Evaluation Pyramid for Non-Deterministic Systems
After the spike, Meridian's evaluation was one thing: 150 golden questions graded by hand by two policy analysts. It was careful and credible. It also took two working days per run. So it ran once a fortnight, prompt changes queued up behind it, and engineers started making "small" changes without waiting. One of those small changes, a rewording meant to make answers friendlier, dropped the condition from seven answers about interest freezes. Nobody noticed for eleven days.
Software engineers know the answer to this shape of problem: the test pyramid. Many fast, cheap tests at the bottom; fewer slow, expensive ones at the top. AI systems need the same shape, but the layers and the graders are different, because many AI outputs have no single right answer that code can compare.
This lesson designs Meridian's evaluation pyramid, chooses graders for each layer, and plans how the datasets stay accurate. It produces part A of MER-08.
Five layers
From the bottom up, each layer is slower, more expensive and closer to real use.
| Layer | What it catches | Size | When it runs | Cost per run |
|---|---|---|---|---|
| 1. Deterministic checks | Wrong figures, missing paragraphs, broken citations, wrong customer | Every output | Every request, in production and tests | Near zero |
| 2. Component evaluations | Retrieval misses, scrubber misses, classifier errors | 400 to 1,000 items each | Every change to that component | A few dollars |
| 3. End-to-end offline | Wrong or incomplete answers, poor summaries and letters | 300 questions, 120 summaries, 100 letters, 150 attacks | Every release candidate | About $40 and 55 minutes |
| 4. Expert review | What the judge misses, new failure types | 50 answers and 30 letters a week | Weekly and before material releases | About 1.5 analyst days a week |
| 5. Online measurement | Real-world drift, adoption, staff feedback | All traffic | Continuously | Dashboard and alert costs |
Layer 1 is special because it runs on every real output, not just in tests. The same code that checks letter figures in the test suite blocks a bad draft in production. It is the layer that enforces the requirements from section 1 that must never be statistical.
Layer 2 measures components on their own. When a policy answer is wrong, you need to know whether retrieval missed the passage or the model misread it. Meridian measures retrieval as recall at six: for 400 questions with known correct sections, how often is the right section among the six chunks sent to the model? Target: 95%.
Layer 3 is the golden set, now grown to 300 questions, graded mostly by a model judge that has been checked against human experts. Layer 4 keeps that judge honest and finds new kinds of failure. Layer 5 watches real use, which section 9 covers in depth.
Choosing the grader
Every check needs a grader, and each kind has a different cost and blind spot.
Code
- Exact, instant, free
- Only for checks with one right answer
- Figures, required text, citation validity, schema
- Use it everywhere it applies
Human expert
- The standard everything else is measured against
- Slow and expensive, and experts disagree sometimes
- Needed for new failure types and for calibration
- Use it for sampling and for checking the judge
Model as judge
- Scales to thousands of items cheaply
- Has known biases that must be measured
- Good for rubric questions with clear criteria
- Use it only after calibration against humans
A model judge has predictable biases. It tends to prefer longer answers, prefer the first of two options it is shown, be lenient, and favour text written by models like itself. Four practices reduce these.
- Ask narrow yes-or-no questions. "Does the answer state every condition in the cited section?" works better than "Rate this answer from 1 to 10."
- Give the judge the evidence. It sees the cited policy passages and the golden answer, not just the output.
- Use a different model family from the one being judged, where possible.
- Calibrate before trusting. Meridian's policy analysts graded 200 answers; the judge graded the same 200. They agreed on 93%. More importantly, the judge passed only 4% of answers the humans failed. That false pass rate is the number that matters, because a judge that is too harsh wastes time while a judge that is too lenient ships bad changes.
The calibration is not a one-off. Every week, analysts re-grade 10% of the items the judge graded. If the false pass rate rises above 5%, the judge is recalibrated before it gates another release.
Datasets that stay true
A golden set decays. Meridian has 20 policy changes a month, and each one can turn a correct golden answer into a wrong one. The design links every golden case to the policy section IDs its answer depends on. When PolicyHub publishes a change, the same policy.published event that re-indexes the document also flags every golden case that cites that section for re-review by an analyst. Without this link, the evaluation would slowly start rewarding outdated answers.
The set also grows from five sources, and each case records where it came from.
- Spike and mystery-shop questions with agreed answers.
- Real staff questions, sampled monthly across products and policy areas.
- Production failures. Every confirmed wrong answer or incident becomes a test case, so the same failure cannot silently return.
- Adversarial cases from the security red team: injections, out-of-scope questions, attempts to extract restricted policy.
- Reviewed variations. A model writes paraphrases of existing questions; an analyst approves each one before it enters the set.
Finally, 20% of the golden set is held out. Prompt engineers never see those cases, and only the release gate runs them. Without a held-out portion, a team slowly tunes its prompts to the visible test cases, and the scores rise faster than real quality.
The evaluation plan is part A of MER-08.
1id: MER-08-A2layers:3 deterministic:4 checks: [figures_match_calculator, mandatory_paragraphs, citations_resolve_current,5 customer_matches_case, schema_valid]6 runs: every output, production and CI7 component:8 retrieval_recall_at_6: {set: 400, target: 0.95}9 scrubber_recall: {set: 500 notes, target: 0.99}10 guard_classifier: {set: 1000, target_recall: 0.95, target_precision: 0.90}11 summary_must_include: {set: 120 cases, target: 1.00 after fallback}12 end_to_end:13 policy_golden: {size: 300, held_out: 60, grader: judge, calibrated: true}14 summaries: {size: 120, grader: judge plus checklist}15 letters: {size: 100, grader: judge plus code}16 adversarial: {size: 150, grader: code plus judge}17 runs: every release candidate18 expert:19 weekly: {answers: 50, letters: 30, judge_recheck: 0.10}20 judge_calibration:21 agreement: 0.9322 false_pass_rate: 0.0423 recalibrate_if_false_pass_above: 0.0524dataset_rules:25 link_cases_to_policy_sections: true26 every_confirmed_incident_becomes_a_case: trueCheck your understanding
0 of 3 answered
1.Why does Meridian measure retrieval recall on its own instead of only scoring final answers?
2.The model judge agrees with human experts on 93% of answers. Which additional number matters most before letting it gate releases?
3.A policy on payment holidays changes. What keeps the golden set from rewarding the old answer?