Enterprise AI Solutions Architecture

Course Content

Enterprise AI Solutions Architecture

13 sections · 29 lessons

From Stakeholder Goal to Measurable Requirement


At the first workshop, Dana Whitfield said her goal was to "make hardship handling faster and safer." Everyone nodded. Two weeks later, the engineering lead asked what "done" would look like, and the room went quiet. Faster than what? Safer in what way? Measured by whom? Nobody had written it down, so every person in the room carried a different picture of success.

A goal is where requirements start, not where they end. A requirement you cannot measure is one you cannot test, cannot defend in review and cannot use to say no to scope creep. The architect's first real piece of work in discovery is turning goals into numbers, with the people who own those numbers agreeing to them.

This lesson builds the first part of MER-03: Meridian's requirements table, with outcome, capability, quality and non-functional requirements, plus the non-goals that keep the scope stable.

What "safer" turned out to meanMake hardshiphandling saferFewer letterswith errorsVulnerablecustomers not missedR-OUT-02:failures 7% to 2%R-LET-01: figureschecked by codeR-SUM-02:recall at least 95%Excerpts shown,human decides
Asking "how would we know?" until every answer is a number with an owner found a requirement nobody had said out loud.

Four layers of requirement

Requirements for an AI system come in four layers. Mixing them up causes most arguments later, so name them separately.

LayerAnswersMeridian exampleOwned by
OutcomeWhat business result should change?Median hardship case handling time falls from 58 to 42 minutesHead of Servicing
CapabilityWhat must the system do?Answer policy questions with cited sectionsArchitect with product owner
QualityHow well must it do it?Cited answers correct, low end of range at least 85%Architect with policy team
Non-functionalUnder what limits?p95 time to first answer text within 2.5 secondsArchitect with engineering

Outcome requirements are special. The system contributes to them but cannot guarantee them, because they also depend on training, adoption and how managers use the saved time. Keep them in the table, because they are the reason the project exists, but do not let anyone treat an outcome as a pass or fail test of the software. The software is tested against capability, quality and non-functional requirements. The outcome is tracked by the business.

From "safer" to something you can count

The technique is simple and slightly annoying: keep asking "how would we know?" until the answer is a number. Watch the work before you ask about it. Sitting with three hardship specialists for a morning taught the architect more than any workshop.

  1. Ask what happens today — "Show me the last hardship letter you wrote, and how long it took."
  2. Ask what bad looks like — "What goes wrong? Who notices? What does it cost?"
  3. Ask how you would know it improved — "What number would change?"
  4. Ask who owns that number — the person who will agree the threshold and track it.
  5. Write it and read it back — the owner signs off the exact wording and threshold.

When the architect asked Dana what "safer" meant, it turned out to mean three different things. Fewer letters with errors. The same policy answer from both servicing centres. And not missing customers who are vulnerable, such as those with serious illness or a recent bereavement. Each became its own requirement with its own measure. The third one, which nobody had mentioned at the workshop, became one of the most important in the design.

A requirement you can test

A testable requirement has six parts: an ID, a statement, a measure, a threshold, a method and an owner. The ID matters more than it looks: it is how the evaluation plan, the threat model and the risk register will refer back to this line months later.

Untestable

  • "The assistant should be accurate"
  • "Answers should be fast"
  • "Letters must be correct"
  • "It must be secure"

Testable

  • R-POL-01: cited answers correct, low end of 95% range at least 85% on the golden set
  • R-NFR-01: p95 time to first answer text 2.5 s or less, measured at the Atlas panel
  • R-LET-01: every figure equals the Hardship Calculator output, checked by code on every draft
  • R-SEC-01: a user never sees a customer or policy outside their entitlements, verified by access tests

Notice that R-POL-01 names where the measure is taken ("on the golden set") and R-NFR-01 names where the clock starts and stops ("at the Atlas panel"). Latency measured at the model provider looks much better than latency the staff member feels, which also includes retrieval, checks and network time.

Non-functional requirements with numbers

Non-functional requirements are where AI systems are most often under-specified. Here are Meridian's, with the reasoning behind each number.

  • Latency. Policy answers: p95 time to first text 2.5 seconds, full answer 9 seconds. Staff are often on a call; a long silence feels broken. Summaries: ready when the case opens, or 12 seconds on demand. Letters: 25 seconds, because the baseline is 22 minutes of manual work.
  • Availability. 99.5% during service hours, 07:00 to 21:00. That allows about 20 minutes of downtime a month in service hours. Staff can always fall back to the manual process, so the assistant must never become a hard dependency.
  • Cost. At most 5 cents per policy answer and 15 cents per letter draft in model and retrieval costs, so the run cost stays inside the business case in section 11.
  • Data. All processing in the bank's region; the provider keeps no prompts and does not train on them.
  • Audit. Every output, its inputs and its model and prompt versions are kept for the records retention period, so any letter can be reconstructed years later.
IDRequirementThresholdMethodOwner
R-OUT-01Hardship case handling time fallsMedian 58 to 42 minutes within 6 monthsAtlas case timestampsHead of Servicing
R-OUT-02Letter quality failures fall7% to 2% or lessExisting QA samplingOps QA lead
R-POL-01Policy answers are correct and citedLow end of range at least 85%300-question golden setPolicy lead
R-POL-02It declines when no policy supports an answerAt least 90% of out-of-scope questions40 out-of-scope casesPolicy lead
R-SUM-01Summary includes every material event100% of must-include itemsCode check against recordsEng lead
R-SUM-02Possible vulnerability mentions are surfacedRecall at least 95%80 labelled case notesVulnerable customers lead
R-LET-01Letter figures equal calculator output100%, every draftCode checkEng lead
R-LET-02Mandatory paragraphs present and unaltered100%, every draftCode checkCompliance
R-LET-03The system cannot send a letterNo send path existsAccess testSecurity
R-NFR-01Policy answer latencyp95 first text 2.5 s, full 9 sPanel telemetryEng lead
R-SEC-01Entitlements are enforcedZero violationsAccess tests per roleSecurity

Non-goals, written under the table and signed by Dana: no customer-facing use, no eligibility or credit decisions, no sending, no writing to core banking, no use in collections calls in the first release.

Check your understanding

0 of 3 answered

1.Why should "hardship case handling time falls from 58 to 42 minutes" not be a pass or fail test of the software?

2.R-NFR-01 says latency is measured "at the Atlas panel". Why state where?

3.Interviews showed "safer" meant three things to Dana. What is the best way to handle that?