Course Content
Enterprise AI Solutions Architecture
13 sections · 29 lessons
Levels of Autonomy and Where Humans Stay in the Loop
Another lender in Meridian's market proudly described its AI document process as "human in the loop." Every AI-drafted customer document was approved by a person before it went out. Then internal audit looked at the logs. The median review took nine seconds for a 400-word document. Over 99% were approved without a single edit. When audit slipped in test documents with obvious errors, reviewers caught fewer than half.
There was a human in the loop. There was no review. The design had put a person at the right point in the process but given them no reason, no time and no tools to actually check anything. Worse, the paperwork said the risk was controlled, so nobody looked further.
This lesson places each Meridian capability on an autonomy scale, chooses where people stay involved, and then designs the review so that it works. It produces part A of MER-06.
An autonomy scale
"Autonomous" and "assistive" are too vague to design with. Meridian uses six levels. Each level is defined by what happens if the AI is wrong and nobody intervenes.
| Level | Name | What the AI does | What the human does | Meridian example |
|---|---|---|---|---|
| L0 | None | Nothing | Everything | Eligibility decisions |
| L1 | Inform | Gives information | Decides and acts | Policy answers, account summary |
| L2 | Draft | Produces an artefact with no effect until approved | Reviews, edits, approves | Hardship letters |
| L3 | Propose action | Prepares a specific system change | Approves each change | Future: setting up a simple plan |
| L4 | Act within limits | Acts alone inside hard limits | Samples afterwards | Candidate: tagging cases by topic |
| L5 | Autonomous | Acts and sets its own limits | Monitors outcomes | Not used in customer-affecting work |
The step from L2 to L3 matters most. At L2 the AI creates an object, such as a draft, that does nothing until a person uses it. At L3 the AI prepares a change to a system of record, and the only thing between the model and the change is the approval. From L3 upward, tool design (the next lesson) becomes as important as review design.
Choosing the level
Four factors decide how much autonomy a capability can have. They are the same factors used for risk surfaces in section 2, plus one more: volume.
- Consequence. What happens to the customer or the bank if the output is wrong?
- Reversibility. Can the effect be undone cheaply and completely?
- Detectability. Will a reviewer, or a later check, notice the error?
- Volume. Can people realistically review every item? Review does not scale to 100,000 items a day.
Regulation adds a fifth consideration. Data protection law gives people protections against decisions based solely on automated processing that significantly affect them, and the EU AI Act requires effective human oversight for systems classified as high-risk. Neither is satisfied by a person who clicks "approve" in nine seconds. What regulators look for is oversight by people with the competence, authority and time to intervene.
For Meridian, the answer was: policy answers and summaries at L1, because staff decide what to do with them; letters at L2, because a letter goes to a customer and cannot be unsent; and nothing at L3 or above in the first release.
Human review that is real
People trust automated output too much when it is usually right. This is called automation bias, and it grows as the system improves: the better the drafts, the less carefully people read them. A review design must work against it.
Review in name only
- A single "Approve" button under a wall of text
- Figures look like any other words
- No idea which parts the AI wrote
- Approval rate and review time never measured
Review that works
- Figures highlighted, each linked to the calculator output
- Mandatory paragraphs shown as locked and already checked
- AI-written paragraphs marked, so attention goes there
- Review time, edits and drill catch rate tracked per team
Meridian's letter review screen does four things. It highlights every figure and shows the calculator value beside it. It shows the mandatory paragraphs as locked, with a tick from the code check, so staff do not waste attention on them. It marks the two or three paragraphs the model wrote, which is where judgement is needed. And it asks one explicit question before approval: "Do the plan terms match what you agreed with the customer?" That question is the one thing no code check can answer.
Then the design measures whether review is happening. The key tool is a review drill: once a month, each reviewer works through 20 drafts in the training environment, four of which contain a seeded error. Drills run only in training, never on real customer letters, so a missed error cannot reach anyone.
| Signal | Healthy | Worrying |
|---|---|---|
| Median letter review time | 2 to 5 minutes | Under 30 seconds |
| Drafts approved with no edits | 30% to 80% | Above 95% for a month |
| Drill catch rate | 90% or more | Under 80% |
| Summary items marked wrong by staff | Stable, low | Zero for months, which suggests nobody is looking |
The last row is subtle. A system with no reported errors for six months is either perfect or unread. Some errors should always be reported; a flat zero is a warning.
Where the human sits
For letters, the full flow shows where people and code each check something.
- Staff choose the plan — the specialist selects the plan the customer agreed on the call.
- Calculator computes — eligibility and every figure come from the rules engine.
- Assistant drafts — the model fills the template and writes the explanation paragraphs.
- Code checks — figures, mandatory paragraphs and customer identity are verified; a failure blocks display.
- Staff review — highlighted figures, locked paragraphs, one explicit confirmation.
- Staff approve — the approval is recorded under the reviewer's own name.
- Existing service sends — DocGen sends the letter; QA samples 5% afterwards, as it does today.
The human is placed where only a human can add value: checking that the letter matches the conversation. Code does the checks that code does better.
The autonomy matrix records each capability's level, the human role and the evidence needed to change the level. It is part A of MER-06.
| Capability | Level | Human role | Review design | Evidence to move up a level |
|---|---|---|---|---|
| Policy answers | L1 | Decides what to tell the customer | Citations with dates | Not planned |
| Account summary | L1 | Uses it to prepare the call | Every line linked to its record | Not planned |
| Letter draft | L2 | Reviews and approves each letter | Highlights, locks, one confirmation | 12 months under 0.5% draft errors and model risk agreement |
| Plan set-up | Not in release 1 | Not applicable | Not applicable | Separate design and validation |
Check your understanding
0 of 3 answered
1.Reviewers approve 99% of drafts unedited in a median of nine seconds. What is the most likely conclusion?
2.What is the key difference between L2 (Draft) and L3 (Propose action)?
3.Why does Meridian's review screen show the mandatory paragraphs as locked and already checked?