Course Content
Enterprise AI Solutions Architecture
13 sections · 29 lessons
From Stakeholder Goal to Measurable Requirement
At the first workshop, Dana Whitfield said her goal was to "make hardship handling faster and safer." Everyone nodded. Two weeks later, the engineering lead asked what "done" would look like, and the room went quiet. Faster than what? Safer in what way? Measured by whom? Nobody had written it down, so every person in the room carried a different picture of success.
A goal is where requirements start, not where they end. A requirement you cannot measure is one you cannot test, cannot defend in review and cannot use to say no to scope creep. The architect's first real piece of work in discovery is turning goals into numbers, with the people who own those numbers agreeing to them.
This lesson builds the first part of MER-03: Meridian's requirements table, with outcome, capability, quality and non-functional requirements, plus the non-goals that keep the scope stable.
Four layers of requirement
Requirements for an AI system come in four layers. Mixing them up causes most arguments later, so name them separately.
| Layer | Answers | Meridian example | Owned by |
|---|---|---|---|
| Outcome | What business result should change? | Median hardship case handling time falls from 58 to 42 minutes | Head of Servicing |
| Capability | What must the system do? | Answer policy questions with cited sections | Architect with product owner |
| Quality | How well must it do it? | Cited answers correct, low end of range at least 85% | Architect with policy team |
| Non-functional | Under what limits? | p95 time to first answer text within 2.5 seconds | Architect with engineering |
Outcome requirements are special. The system contributes to them but cannot guarantee them, because they also depend on training, adoption and how managers use the saved time. Keep them in the table, because they are the reason the project exists, but do not let anyone treat an outcome as a pass or fail test of the software. The software is tested against capability, quality and non-functional requirements. The outcome is tracked by the business.
From "safer" to something you can count
The technique is simple and slightly annoying: keep asking "how would we know?" until the answer is a number. Watch the work before you ask about it. Sitting with three hardship specialists for a morning taught the architect more than any workshop.
- Ask what happens today — "Show me the last hardship letter you wrote, and how long it took."
- Ask what bad looks like — "What goes wrong? Who notices? What does it cost?"
- Ask how you would know it improved — "What number would change?"
- Ask who owns that number — the person who will agree the threshold and track it.
- Write it and read it back — the owner signs off the exact wording and threshold.
When the architect asked Dana what "safer" meant, it turned out to mean three different things. Fewer letters with errors. The same policy answer from both servicing centres. And not missing customers who are vulnerable, such as those with serious illness or a recent bereavement. Each became its own requirement with its own measure. The third one, which nobody had mentioned at the workshop, became one of the most important in the design.
A requirement you can test
A testable requirement has six parts: an ID, a statement, a measure, a threshold, a method and an owner. The ID matters more than it looks: it is how the evaluation plan, the threat model and the risk register will refer back to this line months later.
Untestable
- "The assistant should be accurate"
- "Answers should be fast"
- "Letters must be correct"
- "It must be secure"
Testable
- R-POL-01: cited answers correct, low end of 95% range at least 85% on the golden set
- R-NFR-01: p95 time to first answer text 2.5 s or less, measured at the Atlas panel
- R-LET-01: every figure equals the Hardship Calculator output, checked by code on every draft
- R-SEC-01: a user never sees a customer or policy outside their entitlements, verified by access tests
Notice that R-POL-01 names where the measure is taken ("on the golden set") and R-NFR-01 names where the clock starts and stops ("at the Atlas panel"). Latency measured at the model provider looks much better than latency the staff member feels, which also includes retrieval, checks and network time.
Non-functional requirements with numbers
Non-functional requirements are where AI systems are most often under-specified. Here are Meridian's, with the reasoning behind each number.
- Latency. Policy answers: p95 time to first text 2.5 seconds, full answer 9 seconds. Staff are often on a call; a long silence feels broken. Summaries: ready when the case opens, or 12 seconds on demand. Letters: 25 seconds, because the baseline is 22 minutes of manual work.
- Availability. 99.5% during service hours, 07:00 to 21:00. That allows about 20 minutes of downtime a month in service hours. Staff can always fall back to the manual process, so the assistant must never become a hard dependency.
- Cost. At most 5 cents per policy answer and 15 cents per letter draft in model and retrieval costs, so the run cost stays inside the business case in section 11.
- Data. All processing in the bank's region; the provider keeps no prompts and does not train on them.
- Audit. Every output, its inputs and its model and prompt versions are kept for the records retention period, so any letter can be reconstructed years later.
| ID | Requirement | Threshold | Method | Owner |
|---|---|---|---|---|
| R-OUT-01 | Hardship case handling time falls | Median 58 to 42 minutes within 6 months | Atlas case timestamps | Head of Servicing |
| R-OUT-02 | Letter quality failures fall | 7% to 2% or less | Existing QA sampling | Ops QA lead |
| R-POL-01 | Policy answers are correct and cited | Low end of range at least 85% | 300-question golden set | Policy lead |
| R-POL-02 | It declines when no policy supports an answer | At least 90% of out-of-scope questions | 40 out-of-scope cases | Policy lead |
| R-SUM-01 | Summary includes every material event | 100% of must-include items | Code check against records | Eng lead |
| R-SUM-02 | Possible vulnerability mentions are surfaced | Recall at least 95% | 80 labelled case notes | Vulnerable customers lead |
| R-LET-01 | Letter figures equal calculator output | 100%, every draft | Code check | Eng lead |
| R-LET-02 | Mandatory paragraphs present and unaltered | 100%, every draft | Code check | Compliance |
| R-LET-03 | The system cannot send a letter | No send path exists | Access test | Security |
| R-NFR-01 | Policy answer latency | p95 first text 2.5 s, full 9 s | Panel telemetry | Eng lead |
| R-SEC-01 | Entitlements are enforced | Zero violations | Access tests per role | Security |
Non-goals, written under the table and signed by Dana: no customer-facing use, no eligibility or credit decisions, no sending, no writing to core banking, no use in collections calls in the first release.
Check your understanding
0 of 3 answered
1.Why should "hardship case handling time falls from 58 to 42 minutes" not be a pass or fail test of the software?
2.R-NFR-01 says latency is measured "at the Atlas panel". Why state where?
3.Interviews showed "safer" meant three things to Dana. What is the best way to handle that?