RAG Systems

Course Content

RAG Systems

12 sections · 66 lessons

What is the goal of integrating retrieval with generation?


What you need to know

Putting chunks into a prompt is easy. Getting the model to rely on them, and to show that it did, takes design.

The five integration decisions

  1. Token budget. Decide how many tokens of context to send, for example 3,000. Five 500-token chunks fit; twenty do not.
  2. Deduplication. Overlapping chunks and repeated boilerplate waste the budget. Drop near-duplicates before building the prompt.
  3. Order. Put the strongest passages first. Models use material at the start and end of long prompts more reliably than the middle.
  4. Labels. Number each block and include its source: [2] Leave Policy v4.2, §3.4. The model can then cite [2], and your UI can link it.
  5. Rules. "Answer only from the sources. Cite a source after each sentence. If the sources do not contain the answer, say so."

Check the citations in code

The model can cite a source that does not exist, or write a sentence with no citation. A cheap check after generation catches both:

Python
import redef check_citations(answer: str, sources: dict[int, str]) -> list[str]:    problems = []    for sentence in re.split(r"(?<=[.!?])\s+", answer.strip()):        cited = [int(n) for n in re.findall(r"\[(\d+)\]", sentence)]        if not cited:            problems.append(f"no citation: {sentence!r}")        for n in cited:            if n not in sources:                problems.append(f"cites missing source [{n}]")    return problems

With sources [1] and [2] and an answer whose third sentence ends in [3], it returns ['cites missing source [3]']. The app can then drop that sentence or retry. This does not prove the source supports the sentence; that needs an LLM judge or a quote check, covered in the evaluation section.

Some model APIs now return citations natively: you pass documents as structured inputs and the response marks which spans came from which document. Where available, that is more reliable than parsing [n] from text.

A real-life example

A bank's product-FAQ bot is asked: "Can I close my fixed deposit early, and what is the penalty?"

Retrieval returns three chunks: the premature-withdrawal rule (penalty 1% on the applicable rate), a general FD overview, and a duplicate of the first chunk from an older version of the page. The integration layer drops the duplicate (same text hash), orders the rule first, labels them [1] and [2], and sends 900 tokens of context.

The answer: "Yes, you can close it early [1]. The interest rate is reduced by 1% from the rate applicable for the period the deposit was held [1]." The UI shows [1] as a link to the product terms page. When a customer disputes it, the bank can show the exact source text. Without grounding, the same answer would be a liability.

Follow-up questions to expect

  • "What if two sources contradict each other?" — Prefer the newer or more authoritative one using metadata, and tell the model to mention the conflict rather than silently pick one.
  • "The right chunk was in the prompt and the model still ignored it. Why?" — The model's prior won, or the chunk was buried in noise. Reduce and reorder context, make the rule explicit, and test a stronger model.
  • "How do you show citations to users?" — Link each [n] to the source document at the right page or anchor, using the metadata stored at ingest.