Course Content
Enterprise AI Solutions Architecture
13 sections · 29 lessons
Evidence Gaps and How to Close Them Before You Commit
After the requirements and the decision ladder, Meridian's design looked complete on paper. Retrieval for policy, a workflow for summaries, rules and a template for letters. The evidence register from section 2 told a less comfortable story. Six important claims were still at E0 or E1: vendor slides, public benchmarks and people's beliefs. Some of them, if wrong, would change the architecture.
The temptation at this point is to commit and "find out during build." That is how projects discover in month five that the policy documents cannot be parsed, or that the core banking API is too slow to call during a phone conversation. Finding it out in a three-week spike costs a fraction of finding it out in build.
This lesson turns the evidence register into a spike with pre-agreed decisions, and produces the third part of MER-03: the spike results and the design changes they caused.
Rank the gaps
Not every gap deserves a spike. Rank each one by two questions. If the answer comes back bad, what would change: a parameter, the architecture, or whether the project goes ahead at all? And how cheap is it to test now?
| Gap | If the answer is bad | Cost to test now | Priority |
|---|---|---|---|
| EV-01: answer quality with retrieval | Architecture or project | Low: 150 questions, days | Test now |
| EV-02: policy documents parse cleanly | Architecture of ingestion | Low: 30 documents | Test now |
| EV-03: Ledger API speed | Architecture of the summary | Low: a load test | Test now |
| G-01: regulatory classification of letters | Scope of the letter capability | Medium: legal memo | Start now, runs longer |
| EV-04: provider data terms | Choice of provider | Medium: contract review | Start now, runs longer |
| EV-05: staff trust | Adoption, not architecture | High: needs a working pilot | Test in pilot |
High impact and cheap to test goes first. Staff trust matters a great deal, but it cannot be tested honestly without a working system, and a bad answer would change the rollout plan rather than the architecture. It waits for the pilot.
Designing the spike
Meridian's spike ran for three weeks with three people: the architect, one engineer and half of a senior policy analyst's time. It built a thin slice of the policy answer capability end to end, with no user interface, plus three side tests.
Be clear with everyone about what a spike is not. It is not a prototype that will be polished into production. Its code runs from a laptop, uses exported copies of documents under a data-handling approval, and is thrown away afterwards. What survives is the evidence: the golden set, the scores, the failure analysis and the load test numbers. If a sponsor sees a working spike and asks, "Can we just give this to ten staff next week?", the answer is no, and the reason is everything in sections 4 to 10.
The golden set was the main asset, and it was built in four days, not four months.
- 50 questions from the earlier mystery-shop exercise, which already had agreed answers.
- 70 questions that team leads collected from real staff over two weeks, answered by the policy analyst.
- 30 hard questions written on purpose: answers with conditions, questions about recently changed policies, and out-of-scope questions the assistant should decline.
Two policy analysts graded each answer on the same written guide, without knowing which configuration produced it. They agreed on 91% of cases; a third analyst settled the rest. Measuring that agreement matters: if two experts only agree 70% of the time, the question is badly defined and no model score means much.
Commit, pivot or stop, decided in advance
Before the first run, the sponsor, the architect and the model risk representative signed the criteria below. Writing them first prevents a very human failure: reading whatever result comes back as good enough, because so much has already been invested.
| Test | Commit | Pivot | Stop |
|---|---|---|---|
| Answer correctness, 150 questions | At least 80%, range low end at least 75% | 65% to 80%: fix retrieval before build | Below 65% with good retrieval |
| Policy documents parsed intact | At least 27 of 30 | Fewer: fund a table conversion step | Not applicable |
| Ledger API p95 at 5 requests a second | 300 ms or less | Slower: precompute summaries | Not applicable |
| Legal view on letters | Not high-risk, or manageable | Drafting with extra controls | Letters out of scope |
The "Stop" column is short on purpose. Only two results could end or cut the project: a model that cannot reason over policy even when given the right passage, and a legal view that makes letters impossible. Everything else is a pivot, a change to the design.
What the spike found
The spike scored 125 of 150, or 83%, with a range of about 77% to 88%. That met the commit criteria. The 25 failures mattered more than the score, so the team sorted them by cause.
| Cause | Count | Design change |
|---|---|---|
| Right section not retrieved, usually a table | 11 | Table-aware document conversion |
| Condition dropped from the answer | 6 | Prompt rule plus condition-focused evaluation cases |
| Superseded policy retrieved | 4 | Filter by effective date at retrieval |
| Other | 4 | Tracked, no single fix |
Only 24 of 30 documents parsed with tables intact, so the parsing test hit its pivot, which matched the retrieval failures above. The Ledger API returned p95 of 410 milliseconds at 5 requests a second, also a pivot. And legal's preliminary view was that policy answers and summaries were unlikely to be high-risk, while letters needed a closer look because they touch a decision about a customer's credit. That was manageable: human decision, figures from the calculator, the model only drafting words.
The spike results become part C of MER-03.
1id: MER-03-C2spike: policy answers thin slice, 2026-04-06 to 2026-04-243golden_set: {size: 150, graders: 2, agreement: 0.91}4results:5 answer_correctness: {score: 0.83, range: [0.77, 0.88], verdict: commit}6 parsing_tables_intact: {score: "24/30", verdict: pivot}7 ledger_p95_ms_at_5rps: {score: 410, verdict: pivot}8 legal_letters: {status: preliminary, verdict: proceed with controls}9design_changes:10 - ADR-009 table-aware conversion before chunking11 - ADR-010 effective-date filter in retrieval12 - ADR-011 precompute summaries on hardship.case.opened event13evidence_register_updates:14 EV-01: E3 # was E115 EV-02: E316 EV-03: E317next: grow golden set to 300 before release gate (MER-08)Check your understanding
0 of 3 answered
1.Why did Meridian leave the "staff trust" question for the pilot instead of the spike?
2.Why write the commit, pivot and stop criteria before running the spike?
3.The spike scored 83%, but 11 of 25 failures were tables that were never retrieved. What is the right conclusion?