Course Content
Enterprise AI Solutions Architecture
13 sections · 29 lessons
The Decision Ladder: Rules, Prompt, Retrieval, Agent, Fine-Tune
At the kickoff, two engineers made proposals. The first wanted "an agent with access to every servicing system that handles hardship cases end to end." The second wanted to fine-tune a model on ten years of Meridian's hardship letters so it would "write like us." Both proposals were sincere, and both were technically possible. Both started at the top of a ladder that should be climbed from the bottom.
Every approach to an AI feature sits on a rung. Higher rungs can do more, but each one adds cost, failure modes, evaluation work and validation burden. The decision ladder is a discipline: for each requirement, start at the lowest rung and climb only when evidence shows the rung below cannot meet the requirement.
This lesson walks the five rungs, shows where each of Meridian's capabilities landed and why, and produces the second part of MER-03: the approach decision table.
The five rungs
| Rung | What it is | Good for | What it adds |
|---|---|---|---|
| 1. Rules | Deterministic code, queries, rules engines | Exact values, eligibility, anything with one right answer | Cannot read free text |
| 2. Prompt | A model with instructions and the input | Rewriting, tone, short summaries of supplied text | Variance; knows only its training data |
| 3. Retrieval | Adds your documents at request time | Answers from your own, changing knowledge | An index, freshness, permissions on documents |
| 4. Agent | The model chooses tools and next steps | Open-ended tasks where steps vary by case | Tool permissions, blast radius, hard evaluation |
| 5. Fine-tune | Changes the model's weights with your examples | Narrow format or style at high volume | Training data governance, retraining, re-validation |
Two things are easy to miss. First, the ladder applies per requirement, not per system. One product can use rung 1 for figures, rung 2 for tone and rung 3 for policy, all at once. Second, the rungs are not a quality ranking. Rung 1 is not "less AI" in a bad sense; for an arrears figure, it is the only acceptable rung.
Fine-tuning is at the top, above agents, because of what it adds in a regulated bank. Changing the weights means owning a training data set, a training pipeline, a new model to validate, and a retraining cycle whenever the underlying base model is retired. And it is a poor way to add knowledge that changes: a model fine-tuned on this month's policy is wrong next month, and cannot cite where its answer came from.
Workflow or agent
The most important distinction on the ladder is between rung 2 or 3 used inside a workflow and rung 4, an agent.
Workflow
- Code decides the steps, in a fixed order
- The model fills in one step, such as summarising
- Every run takes the same path, so evaluation is per step
- Failure is local and easy to trace
Agent
- The model decides which tool to call next
- Paths differ from case to case
- Evaluation must cover whole paths, not just outputs
- A wrong early choice can compound over many steps
The test is simple: do you know the steps in advance? For Meridian's account summary, yes. Every summary needs the same data: loan details, payments, arrears, earlier plans, flags and notes. Code can fetch all of it in parallel in about 400 milliseconds, then hand it to a model to summarise. Letting a model decide which records to fetch would add latency, cost and a new way to fail (forgetting to fetch the plan history) for no gain. Agents earn their place when the steps really do vary by case, which section 6 explores.
Climbing only with evidence
Each step up needs evidence that the lower rung fails a specific requirement. "The lower rung might not be good enough" is not evidence. A spike result is.
| Rung | Build time at Meridian | Extra evaluation | Extra validation effort |
|---|---|---|---|
| Rules | Existing calculator, days to integrate | Standard software tests | None new |
| Prompt | 1 to 2 weeks per capability | Golden set and rubric | Light |
| Retrieval | 4 to 6 weeks including parsing | Retrieval metrics plus answers | Moderate |
| Agent | 3 to 5 months with safe tools | Whole-path evaluation, adversarial tests | Heavy |
| Fine-tune | 2 to 4 months plus data work | Everything, again, on each retrain | Heavy, repeated |
These estimates are Meridian's, made by the engineering lead, not universal constants. But the shape is general: the top two rungs cost months, not weeks.
Meridian on the ladder
Policy answers: rung 3, in a workflow. The model does not know Meridian's policies, so a prompt alone fails. Fine-tuning fails too: policies change about 20 times a month, and answers must cite a section. The flow is fixed (rewrite the question, retrieve, answer, check citations), so no agent is needed.
Account summary: rungs 1 and 2, in a workflow. Code fetches the records; a prompt summarises them; code checks the must-include list. Rejected: an agent that "investigates" the account.
Hardship letter: rungs 1 and 2. The Hardship Calculator decides eligibility and computes every figure. The model fills the approved template and writes the explanation paragraphs around fixed figures. Fine-tuning on historical letters was considered and rejected: a sample showed the 7% error rate and wording from withdrawn policies, so training on them would teach the model the bank's old mistakes.
The approach decision table below is part B of MER-03. Each row names the rung chosen, the higher rung rejected, and what would reopen the decision.
| Capability | Rung chosen | Higher rung rejected | Because | Revisit when |
|---|---|---|---|---|
| Policy answers | 3, workflow | Fine-tune | Policy changes monthly; answers need citations | Never for knowledge; maybe for style |
| Account summary | 1 and 2, workflow | Agent | Steps are known in advance | Data needs start to vary by case |
| Letter figures | 1 | Any model | Zero tolerance requirement R-LET-01 | Not planned |
| Letter wording | 2 | Fine-tune | Historical letters contain old errors | Tone rubric fails after prompt work |
| Plan set-up | Not in scope | Agent | Assurance cost far above benefit | After 12 months of stable operation |
| Vulnerability mentions | 2, inside the summary workflow | Fine-tune | Prompt meets R-SUM-02; rules do not | Recall falls below 95% after launch |
Walking one requirement up the ladder
The last row of that table is worth following step by step, because it shows how the two tables above are meant to be used together. The requirement is R-SUM-02: surface possible vulnerability mentions in case notes, with recall of at least 95% on 80 labelled notes. Of those 80 notes, 46 contain a real mention, such as an illness, a bereavement or a job loss.
Start at rung 1, even when you doubt it will work, because the cost table says it takes days. The vulnerable customers team already had a list of 60 terms, such as "chemotherapy", "funeral" and "redundant". Code matched the list against the 80 notes. It found 33 of the 46 mentions, a recall of 72%. The misses were ordinary sentences: "my husband passed in June", "off work since the operation", "things have been hard since the baby came". Adding more terms made little difference, because people describe hard times in endless ways. That is evidence that rung 1 fails this specific requirement, so the team was allowed to climb.
Rung 2 costs one to two weeks, a golden set and light validation. A prompt that asks the model to quote any sentence suggesting possible vulnerability found 44 of the 46 mentions: 96%. It meets the threshold, although 46 cases is a small set, which is one reason recall is measured again after launch.
Now check the rungs above before stopping. Retrieval adds nothing, because code already passes the notes in. An agent adds nothing, because the steps are known. Fine-tuning would cost two to four months, and rung 2 already passes. So the decision stops at rung 2. The keyword list was not thrown away: it still runs as a cheap check, and any note it matches that the model did not surface is logged for the vulnerable customers team. Everything is recorded in one row, including what would reopen it.
Check your understanding
0 of 3 answered
1.Meridian's policies change about 20 times a month and answers must cite a section. Why is fine-tuning the wrong rung for policy knowledge?
2.What is the simplest test for choosing a workflow over an agent?
3.An engineer wants to fine-tune on ten years of hardship letters. What evidence did Meridian find against it?