Enterprise AI Solutions Architecture

Course Content

Enterprise AI Solutions Architecture

13 sections · 29 lessons

The Decision Ladder: Rules, Prompt, Retrieval, Agent, Fine-Tune


At the kickoff, two engineers made proposals. The first wanted "an agent with access to every servicing system that handles hardship cases end to end." The second wanted to fine-tune a model on ten years of Meridian's hardship letters so it would "write like us." Both proposals were sincere, and both were technically possible. Both started at the top of a ladder that should be climbed from the bottom.

Every approach to an AI feature sits on a rung. Higher rungs can do more, but each one adds cost, failure modes, evaluation work and validation burden. The decision ladder is a discipline: for each requirement, start at the lowest rung and climb only when evidence shows the rung below cannot meet the requirement.

This lesson walks the five rungs, shows where each of Meridian's capabilities landed and why, and produces the second part of MER-03: the approach decision table.

Climb only when the rung below fails1 Rules:eligibility, figures2 Prompt: tone,letter wording3 Retrieval:policy answers4 Agent: priced,not in release 15 Fine-tune:rejected, old errorstopbottomEach rung up adds capability, and months of evaluation and validation.
An arrears figure belongs on the bottom rung, and policy that changes 20 times a month can never live in the weights.

The five rungs

RungWhat it isGood forWhat it adds
1. RulesDeterministic code, queries, rules enginesExact values, eligibility, anything with one right answerCannot read free text
2. PromptA model with instructions and the inputRewriting, tone, short summaries of supplied textVariance; knows only its training data
3. RetrievalAdds your documents at request timeAnswers from your own, changing knowledgeAn index, freshness, permissions on documents
4. AgentThe model chooses tools and next stepsOpen-ended tasks where steps vary by caseTool permissions, blast radius, hard evaluation
5. Fine-tuneChanges the model's weights with your examplesNarrow format or style at high volumeTraining data governance, retraining, re-validation

Two things are easy to miss. First, the ladder applies per requirement, not per system. One product can use rung 1 for figures, rung 2 for tone and rung 3 for policy, all at once. Second, the rungs are not a quality ranking. Rung 1 is not "less AI" in a bad sense; for an arrears figure, it is the only acceptable rung.

Fine-tuning is at the top, above agents, because of what it adds in a regulated bank. Changing the weights means owning a training data set, a training pipeline, a new model to validate, and a retraining cycle whenever the underlying base model is retired. And it is a poor way to add knowledge that changes: a model fine-tuned on this month's policy is wrong next month, and cannot cite where its answer came from.

Workflow or agent

The most important distinction on the ladder is between rung 2 or 3 used inside a workflow and rung 4, an agent.

Workflow

  • Code decides the steps, in a fixed order
  • The model fills in one step, such as summarising
  • Every run takes the same path, so evaluation is per step
  • Failure is local and easy to trace

Agent

  • The model decides which tool to call next
  • Paths differ from case to case
  • Evaluation must cover whole paths, not just outputs
  • A wrong early choice can compound over many steps

The test is simple: do you know the steps in advance? For Meridian's account summary, yes. Every summary needs the same data: loan details, payments, arrears, earlier plans, flags and notes. Code can fetch all of it in parallel in about 400 milliseconds, then hand it to a model to summarise. Letting a model decide which records to fetch would add latency, cost and a new way to fail (forgetting to fetch the plan history) for no gain. Agents earn their place when the steps really do vary by case, which section 6 explores.

Climbing only with evidence

Each step up needs evidence that the lower rung fails a specific requirement. "The lower rung might not be good enough" is not evidence. A spike result is.

RungBuild time at MeridianExtra evaluationExtra validation effort
RulesExisting calculator, days to integrateStandard software testsNone new
Prompt1 to 2 weeks per capabilityGolden set and rubricLight
Retrieval4 to 6 weeks including parsingRetrieval metrics plus answersModerate
Agent3 to 5 months with safe toolsWhole-path evaluation, adversarial testsHeavy
Fine-tune2 to 4 months plus data workEverything, again, on each retrainHeavy, repeated

These estimates are Meridian's, made by the engineering lead, not universal constants. But the shape is general: the top two rungs cost months, not weeks.

Meridian on the ladder

Policy answers: rung 3, in a workflow. The model does not know Meridian's policies, so a prompt alone fails. Fine-tuning fails too: policies change about 20 times a month, and answers must cite a section. The flow is fixed (rewrite the question, retrieve, answer, check citations), so no agent is needed.

Account summary: rungs 1 and 2, in a workflow. Code fetches the records; a prompt summarises them; code checks the must-include list. Rejected: an agent that "investigates" the account.

Hardship letter: rungs 1 and 2. The Hardship Calculator decides eligibility and computes every figure. The model fills the approved template and writes the explanation paragraphs around fixed figures. Fine-tuning on historical letters was considered and rejected: a sample showed the 7% error rate and wording from withdrawn policies, so training on them would teach the model the bank's old mistakes.

The approach decision table below is part B of MER-03. Each row names the rung chosen, the higher rung rejected, and what would reopen the decision.

CapabilityRung chosenHigher rung rejectedBecauseRevisit when
Policy answers3, workflowFine-tunePolicy changes monthly; answers need citationsNever for knowledge; maybe for style
Account summary1 and 2, workflowAgentSteps are known in advanceData needs start to vary by case
Letter figures1Any modelZero tolerance requirement R-LET-01Not planned
Letter wording2Fine-tuneHistorical letters contain old errorsTone rubric fails after prompt work
Plan set-upNot in scopeAgentAssurance cost far above benefitAfter 12 months of stable operation
Vulnerability mentions2, inside the summary workflowFine-tunePrompt meets R-SUM-02; rules do notRecall falls below 95% after launch

Walking one requirement up the ladder

The last row of that table is worth following step by step, because it shows how the two tables above are meant to be used together. The requirement is R-SUM-02: surface possible vulnerability mentions in case notes, with recall of at least 95% on 80 labelled notes. Of those 80 notes, 46 contain a real mention, such as an illness, a bereavement or a job loss.

Start at rung 1, even when you doubt it will work, because the cost table says it takes days. The vulnerable customers team already had a list of 60 terms, such as "chemotherapy", "funeral" and "redundant". Code matched the list against the 80 notes. It found 33 of the 46 mentions, a recall of 72%. The misses were ordinary sentences: "my husband passed in June", "off work since the operation", "things have been hard since the baby came". Adding more terms made little difference, because people describe hard times in endless ways. That is evidence that rung 1 fails this specific requirement, so the team was allowed to climb.

Rung 2 costs one to two weeks, a golden set and light validation. A prompt that asks the model to quote any sentence suggesting possible vulnerability found 44 of the 46 mentions: 96%. It meets the threshold, although 46 cases is a small set, which is one reason recall is measured again after launch.

Now check the rungs above before stopping. Retrieval adds nothing, because code already passes the notes in. An agent adds nothing, because the steps are known. Fine-tuning would cost two to four months, and rung 2 already passes. So the decision stops at rung 2. The keyword list was not thrown away: it still runs as a cheap check, and any note it matches that the model did not surface is logged for the vulnerable customers team. Everything is recorded in one row, including what would reopen it.

Check your understanding

0 of 3 answered

1.Meridian's policies change about 20 times a month and answers must cite a section. Why is fine-tuning the wrong rung for policy knowledge?

2.What is the simplest test for choosing a workflow over an agent?

3.An engineer wants to fine-tune on ten years of hardship letters. What evidence did Meridian find against it?