Enterprise AI Solutions Architecture

Course Content

Enterprise AI Solutions Architecture

13 sections · 29 lessons

Designing the Context Pipeline


The context window is the only thing the model knows about a request. It has no access to PolicyHub, no view of the customer's account and no memory of yesterday. Whatever the model uses must be placed in that window by your pipeline, and everything you place there costs money, adds latency and competes for the model's attention.

Meridian's first build of policy answers put the top 25 retrieved chunks into every request, about 18,000 tokens, on the theory that more context meant fewer missed answers. Quality went down, not up. The right passage was usually present, but buried among twenty near-matches from older or unrelated policies, and the model sometimes quoted the wrong one. Cutting to six well-ranked chunks raised correctness, cut time to first text by over a second and cut cost per answer by more than half.

This lesson designs the pipeline that decides what goes into the window, in what order and with what labels. It produces part B of MER-04.

Correct answers by chunks in context, 300 questions80%86%88%84%012325 chunks,6.1 cents6 chunks: chosen3 chunkslose pairsSix well-ranked chunks also cut time to first text by more than a second.
More context adds distraction as well as coverage, so the chunk count is measured on your own documents, not guessed.

The stages of the pipeline

Every capability at Meridian uses the same eight stages. Different capabilities fill them differently, but the stages and their order are fixed, which makes the system easier to test and to explain.

  1. Understand — classify the request and, for policy questions, rewrite it into a clean search query with a small model.
  2. Gather — retrieve policy chunks, fetch records, load the template, in parallel where possible.
  3. Filter — drop anything outside the user's entitlements, outside its effective dates, or for another product.
  4. Rank — rerank candidates, remove near-duplicates, keep the best.
  5. Pack — fit pieces into the token budget by priority, and fail loudly if a required piece does not fit.
  6. Label — wrap each piece with its source ID and mark untrusted content.
  7. Assemble — place pieces in a fixed order: instructions, reference material, case data, question.
  8. Record — write a context manifest listing every piece, version and token count.

The stages are deliberately boring. The interesting decisions are the parameters inside them.

Retrieval choices for policy

The spike showed that most failures were retrieval failures, so this is where the design is most specific.

Chunking. Policies are split on their own section headings, aiming for 300 to 600 tokens per chunk. A table is never split; it stays whole with its caption and the heading above it. This was the table conversion fix from the spike. Each chunk carries metadata: document ID, section number, heading path, effective-from and effective-to dates, product and allowed roles.

Hybrid search. Staff type product codes and internal terms, such as "PL-HR3" for a type of personal loan hardship plan. Embedding search handles meaning well but handles codes badly; keyword search is the opposite. Meridian runs both, merges the results into 40 candidates, then uses a reranking model to choose the best six.

How many chunks. The team measured it rather than guessing, on the 300-question golden set.

Chunks in contextCorrectTime to first text, p95Cost per answer
2580%3.6 s6.1 cents
1286%2.8 s3.4 cents
688%2.3 s2.4 cents
384%2.1 s1.9 cents

Six wins, and three shows that too few also hurts: answers that needed two sections, such as a fee table plus its conditions, started to lose one of them. This curve is specific to Meridian's documents and reranker. The lesson is to measure it, not to copy the number.

A token budget per capability

A budget sets the maximum tokens for each part of the context. It keeps cost predictable, keeps latency inside the requirement and stops one oversized note from pushing out the policy passages.

PartPolicy answerAccount summaryLetter draft
Instructions and rules1,2001,0001,400
Reference material3,600 (6 chunks)300 (must-include list)2,500 (hardship policy)
Case data800 (recent turns)7,000 (records and notes)2,600 (calculator, summary, template)
Question or task200100300
Output limit600700900

The packing code enforces the budget. Required pieces, such as the must-include list or the calculator output, get priority zero and are never silently dropped.

Python
from dataclasses import dataclassclass ContextOverflow(Exception):    pass@dataclassclass Piece:    source_id: str   # e.g. "POL-COL-004#4.2@2026-03-01"    text: str    tokens: int    priority: int    # 0 = required, higher = less important    trusted: booldef pack(pieces: list[Piece], budget: int) -> tuple[list[Piece], list[str]]:    """Keep the most important pieces that fit. Never drop a required one."""    kept, dropped, used = [], [], 0    for piece in sorted(pieces, key=lambda p: p.priority):        if used + piece.tokens <= budget:            kept.append(piece)            used += piece.tokens        elif piece.priority == 0:            raise ContextOverflow(f"required piece does not fit: {piece.source_id}")        else:            dropped.append(piece.source_id)    return kept, droppeddef render(piece: Piece) -> str:    tag = "reference" if piece.trusted else "untrusted_text"    return f'<{tag} source="{piece.source_id}">\n{piece.text}\n</{tag}>'

pack sorts by priority and fills the budget. A required piece that does not fit raises an error instead of producing a summary that quietly lacks the plan history; the orchestrator then trims lower-priority notes and tries again, or shows a clear message. render wraps each piece in a tag that carries its source ID and marks untrusted text.

Labelling, ordering and caching

Labels. Every piece is wrapped with its source and trust level. The instructions tell the model to treat anything inside an untrusted tag as data to read, never as instructions to follow. That alone does not stop prompt injection, and section 10 explains why. But the labels make provenance visible, let output checks confirm that citations point to trusted sources, and make traces readable.

Order. Instructions go first, reference material next, case data after that, and the question last. Models tend to use material at the start and end of a long context more reliably than material buried in the middle, so the question sits at the end, close to where the answer begins.

Caching. Keep the start of the context identical across requests: the same instructions, in the same words, in the same order. Many providers discount repeated prefixes through prompt caching, and they can only do so if the prefix does not change. At Meridian, the 1,200-token instruction block is identical for every policy answer, which section 11 turns into a cost saving.

The context manifest

For every request, the pipeline writes a manifest: which pieces went in, at which versions, how many tokens each used, and what was dropped. It is small, and it is what makes an answer explainable months later.

YAML
request_id: qa-2026-05-12-0931-7f3ccapability: policy_answeruser_role: servicing_agentprompt_version: policy-answer-v14model: gen-a-2026-04-15index_version: policy-idx-2026-05-12T06pieces:  - {source: "POL-HRD-002#3.1@2026-04-01", tokens: 540, priority: 1}  - {source: "POL-HRD-002#3.4@2026-04-01", tokens: 610, priority: 1}  - {source: "POL-COL-004#4.2@2026-03-01", tokens: 480, priority: 2}dropped: ["POL-COL-004#4.5@2026-03-01"]tokens: {input: 5480, output: 312, budget: 6000}

When a team lead asks, "Why did it tell Sam on 12 May that a payment holiday was allowed?", the manifest shows exactly which policy sections the model saw, at which version, and which prompt and model produced the answer.

Section 5 now decides which models receive this context, where they run, and what happens when one of them is unavailable.

Check your understanding

0 of 3 answered

1.Meridian's build put 25 chunks in every request and quality dropped. What is the most likely reason?

2.The account summary's plan history does not fit the token budget because of long notes. What should the pipeline do?

3.Why does Meridian keep the 1,200-token instruction block identical, word for word, on every policy answer?