Course Content
RAG Systems
12 sections · 66 lessons
What is the goal of integrating retrieval with generation?
What you need to know
Putting chunks into a prompt is easy. Getting the model to rely on them, and to show that it did, takes design.
The five integration decisions
- Token budget. Decide how many tokens of context to send, for example 3,000. Five 500-token chunks fit; twenty do not.
- Deduplication. Overlapping chunks and repeated boilerplate waste the budget. Drop near-duplicates before building the prompt.
- Order. Put the strongest passages first. Models use material at the start and end of long prompts more reliably than the middle.
- Labels. Number each block and include its source:
[2] Leave Policy v4.2, §3.4. The model can then cite[2], and your UI can link it. - Rules. "Answer only from the sources. Cite a source after each sentence. If the sources do not contain the answer, say so."
Check the citations in code
The model can cite a source that does not exist, or write a sentence with no citation. A cheap check after generation catches both:
1import re23def check_citations(answer: str, sources: dict[int, str]) -> list[str]:4 problems = []5 for sentence in re.split(r"(?<=[.!?])\s+", answer.strip()):6 cited = [int(n) for n in re.findall(r"\[(\d+)\]", sentence)]7 if not cited:8 problems.append(f"no citation: {sentence!r}")9 for n in cited:10 if n not in sources:11 problems.append(f"cites missing source [{n}]")12 return problemsWith sources [1] and [2] and an answer whose third sentence ends in [3], it returns ['cites missing source [3]']. The app can then drop that sentence or retry. This does not prove the source supports the sentence; that needs an LLM judge or a quote check, covered in the evaluation section.
Some model APIs now return citations natively: you pass documents as structured inputs and the response marks which spans came from which document. Where available, that is more reliable than parsing [n] from text.
A real-life example
A bank's product-FAQ bot is asked: "Can I close my fixed deposit early, and what is the penalty?"
Retrieval returns three chunks: the premature-withdrawal rule (penalty 1% on the applicable rate), a general FD overview, and a duplicate of the first chunk from an older version of the page. The integration layer drops the duplicate (same text hash), orders the rule first, labels them [1] and [2], and sends 900 tokens of context.
The answer: "Yes, you can close it early [1]. The interest rate is reduced by 1% from the rate applicable for the period the deposit was held [1]." The UI shows [1] as a link to the product terms page. When a customer disputes it, the bank can show the exact source text. Without grounding, the same answer would be a liability.
Follow-up questions to expect
- "What if two sources contradict each other?" — Prefer the newer or more authoritative one using metadata, and tell the model to mention the conflict rather than silently pick one.
- "The right chunk was in the prompt and the model still ignored it. Why?" — The model's prior won, or the chunk was buried in noise. Reduce and reorder context, make the rule explicit, and test a stronger model.
- "How do you show citations to users?" — Link each
[n]to the source document at the right page or anchor, using the metadata stored at ingest.