Course Content
Enterprise AI Solutions Architecture
13 sections · 29 lessons
The Building Blocks of AI, and Where Each One Breaks
"Summarise the customer's account history" sounds like one capability. At Meridian it is at least four. The system must fetch loan and payment records, pull events such as promises to pay out of free-text case notes, notice any mention of illness or job loss, and then write a short, readable summary. Each of those steps fails in a different way, and each needs a different control.
Vendors sell whole products: "a servicing copilot", "an AI letter writer". Architects need to see the parts, because risk lives in the parts. A summary that invents a payment arrangement and a summary that leaves out a previous hardship plan are both "bad summaries", but the first is visible to a careful reader and the second is not.
This lesson gives you a small vocabulary of capability primitives, the typical failure of each, and a way to map them onto a use case. The map becomes the first part of MER-02, the capability and evidence register.
Eight primitives
Almost every enterprise AI feature is built from a handful of primitives. Learn their failure modes once and you can reason about any new product.
| Primitive | What it does | Typical failure |
|---|---|---|
| Generation | Writes new text from instructions | Invents facts, wrong tone, makes promises |
| Summarisation | Shortens source material | Leaves out important facts, shifts emphasis |
| Extraction | Pulls fields out of messy text | Wrong value, missed value, invented value |
| Classification | Picks a label from a set | Wrong label, overconfident, labels drift over time |
| Semantic retrieval | Finds passages by meaning | Misses the right passage, returns an outdated one |
| Grounded question answering | Answers from supplied passages | Drops a condition, reasons past the evidence |
| Tool use | Chooses a function and its arguments | Wrong tool, wrong argument, unsafe action |
| Transformation | Rewrites, translates, fills templates | Meaning changes quietly |
Two notes. First, some steps in a feature are not AI at all. Fetching the loan balance is an API call, and it should stay one. A common design error is to let a model do work that a query does better. Second, the primitives combine. Grounded question answering almost always depends on retrieval, so a retrieval failure shows up as a wrong answer, and you will look in the wrong place unless you measure each primitive on its own.
What makes a risk surface
A primitive's risk surface is the set of ways it can cause harm in one specific context. The same primitive can be low risk in one place and high risk in another. Four factors decide it.
- Consequence. What happens if the output is wrong? A clumsy sentence in a draft is cheap. A wrong statement about eligibility can mislead a customer in financial difficulty.
- Reversibility. Can the harm be undone? An answer shown to a staff member can be corrected. A letter sent to a customer cannot be unsent.
- Detectability. Will the person reviewing the output notice the error? This is the factor teams most often forget.
- Exposure to untrusted input. Does the primitive read text that an outsider wrote? Case notes at Meridian include pasted customer emails, and anyone can write anything in an email.
Detectability deserves its own paragraph. A reviewer can check what is on the screen, but cannot easily check what is missing from it. An invented payment arrangement sits in the summary, where a careful agent may spot it. A missing previous hardship plan leaves no trace at all. So omission is often a more dangerous failure than invention, even though invention gets more attention.
Errors a reviewer tends to catch
- A figure that looks wrong
- A sentence that contradicts the source link
- An odd or unkind tone
- An answer to a different question
Errors a reviewer tends to miss
- A condition dropped from a policy answer
- A previous hardship plan left out of a summary
- An outdated policy cited with confidence
- A vulnerability disclosure buried in old notes
Decomposing Meridian's three capabilities
Here is how the architect broke Meridian's three capabilities into primitives. The "not AI" steps are listed on purpose, so nobody later moves them into a prompt.
| Capability | Not AI | AI primitives |
|---|---|---|
| Policy answers | Staff identity and entitlements | Transformation (rewrite the question), semantic retrieval, grounded question answering |
| Account summary | Fetch loans, payments, arrears, flags | Extraction from notes, classification of possible vulnerability, summarisation |
| Hardship letter | Eligibility, plan terms, every figure | Transformation (fill the template), generation (explanation paragraphs) |
The account summary raised a hard question. Case notes sometimes mention an illness or a bereavement that was never recorded as a formal vulnerability flag. Should the assistant detect these and set the flag? The architect's answer was no. Health information is special-category personal data under data protection law, and inferring it automatically creates both privacy and fairness risks. Instead, the summary shows the exact note excerpt, marked "possible vulnerability, please check", and a human decides. The model surfaces; people decide.
The capability and risk map
The map lists each primitive in use, its failure in this context, and the four factors. The last column is a first idea for a control, not a final design; sections 6 and 10 turn these ideas into specified controls.
| Primitive in use | Failure here | Consequence | Detectability | Untrusted input | First control idea |
|---|---|---|---|---|---|
| Retrieval of policy | Superseded policy returned | High | Low | Low | Filter by effective date; show dates |
| Grounded answering | Condition dropped | High | Low | Low | Evaluation cases built around conditions |
| Extraction from notes | Invented promise to pay | High | Medium | High | Every line links to its source note |
| Summarisation | Earlier hardship plan omitted | High | Low | High | Must-include list computed from records |
| Vulnerability classification | Disclosure missed | High | Low | High | Show raw excerpts; never clear a flag |
| Letter generation | Confusing or unkind wording | Medium | High | Low | Approved template, tone rubric, review |
| Tool use for data fetch | Wrong customer's data | High | Low | Medium | Orchestrator binds the customer ID |
Read down the detectability column. Five of seven are "Low". That single column explains most of the design you will see in later sections: citations with dates, links to source records, a must-include checklist computed by code, and the rule that the customer ID comes from the case, never from the model.
The map above is part A of MER-02. The next lesson adds the second part: what the bank actually knows about each capability, and how strong that knowledge is.
Check your understanding
0 of 3 answered
1.Why is a missing earlier hardship plan in a summary often more dangerous than an invented one?
2.A case note contains a customer's email saying she has just started cancer treatment, but no vulnerability flag is set. What did Meridian's design choose?
3.Which step in the account summary should not be an AI primitive at all?