Enterprise AI Solutions Architecture

Course Content

Enterprise AI Solutions Architecture

13 sections · 29 lessons

Evidence Gaps and How to Close Them Before You Commit


After the requirements and the decision ladder, Meridian's design looked complete on paper. Retrieval for policy, a workflow for summaries, rules and a template for letters. The evidence register from section 2 told a less comfortable story. Six important claims were still at E0 or E1: vendor slides, public benchmarks and people's beliefs. Some of them, if wrong, would change the architecture.

The temptation at this point is to commit and "find out during build." That is how projects discover in month five that the policy documents cannot be parsed, or that the core banking API is too slow to call during a phone conversation. Finding it out in a three-week spike costs a fraction of finding it out in build.

This lesson turns the evidence register into a spike with pre-agreed decisions, and produces the third part of MER-03: the spike results and the design changes they caused.

Three weeks of spike, four verdicts83%, range77 to 88commit24 of 30pivot410 ms at p95pivotmanageableproceedSpike resultVerdictAnswer correctnessTables parsed intactLedger API speedLegal view on lettersCriteria were signed before the first run.
The spike met its commit bar and still forced two pivots, found in days instead of months, which is exactly its job.

Rank the gaps

Not every gap deserves a spike. Rank each one by two questions. If the answer comes back bad, what would change: a parameter, the architecture, or whether the project goes ahead at all? And how cheap is it to test now?

GapIf the answer is badCost to test nowPriority
EV-01: answer quality with retrievalArchitecture or projectLow: 150 questions, daysTest now
EV-02: policy documents parse cleanlyArchitecture of ingestionLow: 30 documentsTest now
EV-03: Ledger API speedArchitecture of the summaryLow: a load testTest now
G-01: regulatory classification of lettersScope of the letter capabilityMedium: legal memoStart now, runs longer
EV-04: provider data termsChoice of providerMedium: contract reviewStart now, runs longer
EV-05: staff trustAdoption, not architectureHigh: needs a working pilotTest in pilot

High impact and cheap to test goes first. Staff trust matters a great deal, but it cannot be tested honestly without a working system, and a bad answer would change the rollout plan rather than the architecture. It waits for the pilot.

Designing the spike

Meridian's spike ran for three weeks with three people: the architect, one engineer and half of a senior policy analyst's time. It built a thin slice of the policy answer capability end to end, with no user interface, plus three side tests.

Be clear with everyone about what a spike is not. It is not a prototype that will be polished into production. Its code runs from a laptop, uses exported copies of documents under a data-handling approval, and is thrown away afterwards. What survives is the evidence: the golden set, the scores, the failure analysis and the load test numbers. If a sponsor sees a working spike and asks, "Can we just give this to ten staff next week?", the answer is no, and the reason is everything in sections 4 to 10.

The golden set was the main asset, and it was built in four days, not four months.

  • 50 questions from the earlier mystery-shop exercise, which already had agreed answers.
  • 70 questions that team leads collected from real staff over two weeks, answered by the policy analyst.
  • 30 hard questions written on purpose: answers with conditions, questions about recently changed policies, and out-of-scope questions the assistant should decline.

Two policy analysts graded each answer on the same written guide, without knowing which configuration produced it. They agreed on 91% of cases; a third analyst settled the rest. Measuring that agreement matters: if two experts only agree 70% of the time, the question is badly defined and no model score means much.

Commit, pivot or stop, decided in advance

Before the first run, the sponsor, the architect and the model risk representative signed the criteria below. Writing them first prevents a very human failure: reading whatever result comes back as good enough, because so much has already been invested.

TestCommitPivotStop
Answer correctness, 150 questionsAt least 80%, range low end at least 75%65% to 80%: fix retrieval before buildBelow 65% with good retrieval
Policy documents parsed intactAt least 27 of 30Fewer: fund a table conversion stepNot applicable
Ledger API p95 at 5 requests a second300 ms or lessSlower: precompute summariesNot applicable
Legal view on lettersNot high-risk, or manageableDrafting with extra controlsLetters out of scope

The "Stop" column is short on purpose. Only two results could end or cut the project: a model that cannot reason over policy even when given the right passage, and a legal view that makes letters impossible. Everything else is a pivot, a change to the design.

What the spike found

The spike scored 125 of 150, or 83%, with a range of about 77% to 88%. That met the commit criteria. The 25 failures mattered more than the score, so the team sorted them by cause.

CauseCountDesign change
Right section not retrieved, usually a table11Table-aware document conversion
Condition dropped from the answer6Prompt rule plus condition-focused evaluation cases
Superseded policy retrieved4Filter by effective date at retrieval
Other4Tracked, no single fix

Only 24 of 30 documents parsed with tables intact, so the parsing test hit its pivot, which matched the retrieval failures above. The Ledger API returned p95 of 410 milliseconds at 5 requests a second, also a pivot. And legal's preliminary view was that policy answers and summaries were unlikely to be high-risk, while letters needed a closer look because they touch a decision about a customer's credit. That was manageable: human decision, figures from the calculator, the model only drafting words.

The spike results become part C of MER-03.

YAML
id: MER-03-Cspike: policy answers thin slice, 2026-04-06 to 2026-04-24golden_set: {size: 150, graders: 2, agreement: 0.91}results:  answer_correctness: {score: 0.83, range: [0.77, 0.88], verdict: commit}  parsing_tables_intact: {score: "24/30", verdict: pivot}  ledger_p95_ms_at_5rps: {score: 410, verdict: pivot}  legal_letters: {status: preliminary, verdict: proceed with controls}design_changes:  - ADR-009 table-aware conversion before chunking  - ADR-010 effective-date filter in retrieval  - ADR-011 precompute summaries on hardship.case.opened eventevidence_register_updates:  EV-01: E3   # was E1  EV-02: E3  EV-03: E3next: grow golden set to 300 before release gate (MER-08)

Check your understanding

0 of 3 answered

1.Why did Meridian leave the "staff trust" question for the pilot instead of the spike?

2.Why write the commit, pivot and stop criteria before running the spike?

3.The spike scored 83%, but 11 of 25 failures were tables that were never retrieved. What is the right conclusion?