Enterprise AI Solutions Architecture

Course Content

Enterprise AI Solutions Architecture

13 sections · 29 lessons

Levels of Autonomy and Where Humans Stay in the Loop


Another lender in Meridian's market proudly described its AI document process as "human in the loop." Every AI-drafted customer document was approved by a person before it went out. Then internal audit looked at the logs. The median review took nine seconds for a 400-word document. Over 99% were approved without a single edit. When audit slipped in test documents with obvious errors, reviewers caught fewer than half.

There was a human in the loop. There was no review. The design had put a person at the right point in the process but given them no reason, no time and no tools to actually check anything. Worse, the paperwork said the risk was controlled, so nobody looked further.

This lesson places each Meridian capability on an autonomy scale, chooses where people stay involved, and then designs the review so that it works. It produces part A of MER-06.

Letter flow: who checks whatStaffchoose the planCalculatorcomputesModeldrafts wordingCode checksthe draftStaffreview, approveDocGen sendsblocksbad draftsdrills:92% caughtThe human sits where only a human adds value: does it match the call?
A person in the loop is a control only if the review screen gives them a reason, time and tools to check.

An autonomy scale

"Autonomous" and "assistive" are too vague to design with. Meridian uses six levels. Each level is defined by what happens if the AI is wrong and nobody intervenes.

LevelNameWhat the AI doesWhat the human doesMeridian example
L0NoneNothingEverythingEligibility decisions
L1InformGives informationDecides and actsPolicy answers, account summary
L2DraftProduces an artefact with no effect until approvedReviews, edits, approvesHardship letters
L3Propose actionPrepares a specific system changeApproves each changeFuture: setting up a simple plan
L4Act within limitsActs alone inside hard limitsSamples afterwardsCandidate: tagging cases by topic
L5AutonomousActs and sets its own limitsMonitors outcomesNot used in customer-affecting work

The step from L2 to L3 matters most. At L2 the AI creates an object, such as a draft, that does nothing until a person uses it. At L3 the AI prepares a change to a system of record, and the only thing between the model and the change is the approval. From L3 upward, tool design (the next lesson) becomes as important as review design.

Choosing the level

Four factors decide how much autonomy a capability can have. They are the same factors used for risk surfaces in section 2, plus one more: volume.

  • Consequence. What happens to the customer or the bank if the output is wrong?
  • Reversibility. Can the effect be undone cheaply and completely?
  • Detectability. Will a reviewer, or a later check, notice the error?
  • Volume. Can people realistically review every item? Review does not scale to 100,000 items a day.

Regulation adds a fifth consideration. Data protection law gives people protections against decisions based solely on automated processing that significantly affect them, and the EU AI Act requires effective human oversight for systems classified as high-risk. Neither is satisfied by a person who clicks "approve" in nine seconds. What regulators look for is oversight by people with the competence, authority and time to intervene.

For Meridian, the answer was: policy answers and summaries at L1, because staff decide what to do with them; letters at L2, because a letter goes to a customer and cannot be unsent; and nothing at L3 or above in the first release.

Human review that is real

People trust automated output too much when it is usually right. This is called automation bias, and it grows as the system improves: the better the drafts, the less carefully people read them. A review design must work against it.

Review in name only

  • A single "Approve" button under a wall of text
  • Figures look like any other words
  • No idea which parts the AI wrote
  • Approval rate and review time never measured

Review that works

  • Figures highlighted, each linked to the calculator output
  • Mandatory paragraphs shown as locked and already checked
  • AI-written paragraphs marked, so attention goes there
  • Review time, edits and drill catch rate tracked per team

Meridian's letter review screen does four things. It highlights every figure and shows the calculator value beside it. It shows the mandatory paragraphs as locked, with a tick from the code check, so staff do not waste attention on them. It marks the two or three paragraphs the model wrote, which is where judgement is needed. And it asks one explicit question before approval: "Do the plan terms match what you agreed with the customer?" That question is the one thing no code check can answer.

Then the design measures whether review is happening. The key tool is a review drill: once a month, each reviewer works through 20 drafts in the training environment, four of which contain a seeded error. Drills run only in training, never on real customer letters, so a missed error cannot reach anyone.

SignalHealthyWorrying
Median letter review time2 to 5 minutesUnder 30 seconds
Drafts approved with no edits30% to 80%Above 95% for a month
Drill catch rate90% or moreUnder 80%
Summary items marked wrong by staffStable, lowZero for months, which suggests nobody is looking

The last row is subtle. A system with no reported errors for six months is either perfect or unread. Some errors should always be reported; a flat zero is a warning.

Where the human sits

For letters, the full flow shows where people and code each check something.

  1. Staff choose the plan — the specialist selects the plan the customer agreed on the call.
  2. Calculator computes — eligibility and every figure come from the rules engine.
  3. Assistant drafts — the model fills the template and writes the explanation paragraphs.
  4. Code checks — figures, mandatory paragraphs and customer identity are verified; a failure blocks display.
  5. Staff review — highlighted figures, locked paragraphs, one explicit confirmation.
  6. Staff approve — the approval is recorded under the reviewer's own name.
  7. Existing service sends — DocGen sends the letter; QA samples 5% afterwards, as it does today.

The human is placed where only a human can add value: checking that the letter matches the conversation. Code does the checks that code does better.

The autonomy matrix records each capability's level, the human role and the evidence needed to change the level. It is part A of MER-06.

CapabilityLevelHuman roleReview designEvidence to move up a level
Policy answersL1Decides what to tell the customerCitations with datesNot planned
Account summaryL1Uses it to prepare the callEvery line linked to its recordNot planned
Letter draftL2Reviews and approves each letterHighlights, locks, one confirmation12 months under 0.5% draft errors and model risk agreement
Plan set-upNot in release 1Not applicableNot applicableSeparate design and validation

Check your understanding

0 of 3 answered

1.Reviewers approve 99% of drafts unedited in a median of nine seconds. What is the most likely conclusion?

2.What is the key difference between L2 (Draft) and L3 (Propose action)?

3.Why does Meridian's review screen show the mandatory paragraphs as locked and already checked?