Course Content
Enterprise AI Solutions Architecture
13 sections · 29 lessons
Designing the Context Pipeline
The context window is the only thing the model knows about a request. It has no access to PolicyHub, no view of the customer's account and no memory of yesterday. Whatever the model uses must be placed in that window by your pipeline, and everything you place there costs money, adds latency and competes for the model's attention.
Meridian's first build of policy answers put the top 25 retrieved chunks into every request, about 18,000 tokens, on the theory that more context meant fewer missed answers. Quality went down, not up. The right passage was usually present, but buried among twenty near-matches from older or unrelated policies, and the model sometimes quoted the wrong one. Cutting to six well-ranked chunks raised correctness, cut time to first text by over a second and cut cost per answer by more than half.
This lesson designs the pipeline that decides what goes into the window, in what order and with what labels. It produces part B of MER-04.
The stages of the pipeline
Every capability at Meridian uses the same eight stages. Different capabilities fill them differently, but the stages and their order are fixed, which makes the system easier to test and to explain.
- Understand — classify the request and, for policy questions, rewrite it into a clean search query with a small model.
- Gather — retrieve policy chunks, fetch records, load the template, in parallel where possible.
- Filter — drop anything outside the user's entitlements, outside its effective dates, or for another product.
- Rank — rerank candidates, remove near-duplicates, keep the best.
- Pack — fit pieces into the token budget by priority, and fail loudly if a required piece does not fit.
- Label — wrap each piece with its source ID and mark untrusted content.
- Assemble — place pieces in a fixed order: instructions, reference material, case data, question.
- Record — write a context manifest listing every piece, version and token count.
The stages are deliberately boring. The interesting decisions are the parameters inside them.
Retrieval choices for policy
The spike showed that most failures were retrieval failures, so this is where the design is most specific.
Chunking. Policies are split on their own section headings, aiming for 300 to 600 tokens per chunk. A table is never split; it stays whole with its caption and the heading above it. This was the table conversion fix from the spike. Each chunk carries metadata: document ID, section number, heading path, effective-from and effective-to dates, product and allowed roles.
Hybrid search. Staff type product codes and internal terms, such as "PL-HR3" for a type of personal loan hardship plan. Embedding search handles meaning well but handles codes badly; keyword search is the opposite. Meridian runs both, merges the results into 40 candidates, then uses a reranking model to choose the best six.
How many chunks. The team measured it rather than guessing, on the 300-question golden set.
| Chunks in context | Correct | Time to first text, p95 | Cost per answer |
|---|---|---|---|
| 25 | 80% | 3.6 s | 6.1 cents |
| 12 | 86% | 2.8 s | 3.4 cents |
| 6 | 88% | 2.3 s | 2.4 cents |
| 3 | 84% | 2.1 s | 1.9 cents |
Six wins, and three shows that too few also hurts: answers that needed two sections, such as a fee table plus its conditions, started to lose one of them. This curve is specific to Meridian's documents and reranker. The lesson is to measure it, not to copy the number.
A token budget per capability
A budget sets the maximum tokens for each part of the context. It keeps cost predictable, keeps latency inside the requirement and stops one oversized note from pushing out the policy passages.
| Part | Policy answer | Account summary | Letter draft |
|---|---|---|---|
| Instructions and rules | 1,200 | 1,000 | 1,400 |
| Reference material | 3,600 (6 chunks) | 300 (must-include list) | 2,500 (hardship policy) |
| Case data | 800 (recent turns) | 7,000 (records and notes) | 2,600 (calculator, summary, template) |
| Question or task | 200 | 100 | 300 |
| Output limit | 600 | 700 | 900 |
The packing code enforces the budget. Required pieces, such as the must-include list or the calculator output, get priority zero and are never silently dropped.
1from dataclasses import dataclass23class ContextOverflow(Exception):4 pass56@dataclass7class Piece:8 source_id: str # e.g. "POL-COL-004#4.2@2026-03-01"9 text: str10 tokens: int11 priority: int # 0 = required, higher = less important12 trusted: bool1314def pack(pieces: list[Piece], budget: int) -> tuple[list[Piece], list[str]]:15 """Keep the most important pieces that fit. Never drop a required one."""16 kept, dropped, used = [], [], 017 for piece in sorted(pieces, key=lambda p: p.priority):18 if used + piece.tokens <= budget:19 kept.append(piece)20 used += piece.tokens21 elif piece.priority == 0:22 raise ContextOverflow(f"required piece does not fit: {piece.source_id}")23 else:24 dropped.append(piece.source_id)25 return kept, dropped2627def render(piece: Piece) -> str:28 tag = "reference" if piece.trusted else "untrusted_text"29 return f'<{tag} source="{piece.source_id}">\n{piece.text}\n</{tag}>'pack sorts by priority and fills the budget. A required piece that does not fit raises an error instead of producing a summary that quietly lacks the plan history; the orchestrator then trims lower-priority notes and tries again, or shows a clear message. render wraps each piece in a tag that carries its source ID and marks untrusted text.
Labelling, ordering and caching
Labels. Every piece is wrapped with its source and trust level. The instructions tell the model to treat anything inside an untrusted tag as data to read, never as instructions to follow. That alone does not stop prompt injection, and section 10 explains why. But the labels make provenance visible, let output checks confirm that citations point to trusted sources, and make traces readable.
Order. Instructions go first, reference material next, case data after that, and the question last. Models tend to use material at the start and end of a long context more reliably than material buried in the middle, so the question sits at the end, close to where the answer begins.
Caching. Keep the start of the context identical across requests: the same instructions, in the same words, in the same order. Many providers discount repeated prefixes through prompt caching, and they can only do so if the prefix does not change. At Meridian, the 1,200-token instruction block is identical for every policy answer, which section 11 turns into a cost saving.
The context manifest
For every request, the pipeline writes a manifest: which pieces went in, at which versions, how many tokens each used, and what was dropped. It is small, and it is what makes an answer explainable months later.
1request_id: qa-2026-05-12-0931-7f3c2capability: policy_answer3user_role: servicing_agent4prompt_version: policy-answer-v145model: gen-a-2026-04-156index_version: policy-idx-2026-05-12T067pieces:8 - {source: "POL-HRD-002#3.1@2026-04-01", tokens: 540, priority: 1}9 - {source: "POL-HRD-002#3.4@2026-04-01", tokens: 610, priority: 1}10 - {source: "POL-COL-004#4.2@2026-03-01", tokens: 480, priority: 2}11dropped: ["POL-COL-004#4.5@2026-03-01"]12tokens: {input: 5480, output: 312, budget: 6000}When a team lead asks, "Why did it tell Sam on 12 May that a payment holiday was allowed?", the manifest shows exactly which policy sections the model saw, at which version, and which prompt and model produced the answer.
Section 5 now decides which models receive this context, where they run, and what happens when one of them is unavailable.
Check your understanding
0 of 3 answered
1.Meridian's build put 25 chunks in every request and quality dropped. What is the most likely reason?
2.The account summary's plan history does not fit the token budget because of long notes. What should the pipeline do?
3.Why does Meridian keep the 1,200-token instruction block identical, word for word, on every policy answer?