Course Content
RAG Systems
12 sections · 66 lessons
How do you limit the LLM to answer only from retrieved context?
What you need to know
A model has two sources of knowledge: the context you give it and what it learned in training (called parametric knowledge). You cannot switch the second off. You can make the first the easiest path, and catch the cases where the model leaves it.
- Instruction and escape hatch — "Use only the documents. If they do not answer the question, reply exactly: …"
- Retrieval gate — if zero chunks pass the filters, or the best reranker score is below a threshold, return the fallback without calling the model.
- Mandatory citations — every sentence ends with a chunk id such as
[2]. - Citation check in code — every cited id was retrieved; every sentence has at least one.
- Faithfulness check — a judge model checks whether each claim is supported by its cited chunk, on all traffic for high-risk products or a sample for others.
- Low temperature — where the model allows it, 0 to 0.2 so it does not wander.
The citation check (step 4)
This is cheap, deterministic and runs on every request.
1import re23REFUSAL = "I could not find that in the guidelines."45def check_answer(answer: str, retrieved_ids: set[int]) -> list[str]:6 """Return a list of problems; an empty list means the answer passes."""7 if answer.strip() == REFUSAL:8 return []9 problems = []10 for s in re.split(r"(?<=[.!?])\s+", answer.strip()):11 cited = {int(n) for n in re.findall(r"\[(\d+)\]", s)}12 if not cited:13 problems.append(f"no citation: {s[:60]}")14 elif not cited <= retrieved_ids:15 problems.append(f"cites unknown source {sorted(cited - retrieved_ids)}")16 return problems1718ans = ("Give the first antibiotic dose within one hour [1]. "19 "Repeat lactate after 2 hours [4]. Most patients recover fully.")20print(check_answer(ans, retrieved_ids={1, 2, 3}))21# ['cites unknown source [4]', 'no citation: Most patients recover fully.']The function splits the answer into sentences, pulls out [n] markers, and flags sentences with no citation or with a citation to a chunk that was never retrieved. An invented citation is a strong sign of an invented claim. When it fails, you can retry once with the problems listed, remove the uncited sentence, or return the fallback.
Some model APIs now return citations natively, with the exact quoted span from each document. When available, they make step 4 simpler and more reliable than parsing [n] markers.
The trade-off to name
Strict grounding also blocks harmless reasoning, such as adding two numbers from the context or answering "yes" when the document implies it. Measure two numbers on your evaluation set: the faithfulness score (share of claims supported) and the unhelpful-refusal rate (share of answerable questions the system refused). Tune the gate threshold and prompt between them.
A real-life example
A hospital's sepsis-protocol assistant must never give advice that is not in the approved protocol. The team runs 400 real staff questions through three versions:
- Prompt only: many answers include general medical advice not in the protocol, usually correct, but unapproved.
- Prompt and retrieval gate: the unapproved advice drops sharply, because questions with no good match now get the fallback.
- Prompt, gate and citation check with one retry: almost every remaining uncited sentence is caught; the few left are flagged for the weekly review.
The refusal rate on answerable questions rises from 4% to 9%. The clinical lead accepts that, because a refusal sends a nurse to the full protocol, while an unapproved claim could harm a patient. For the bank's FAQ bot on the same platform, the team sets a looser threshold, because an unnecessary refusal there only costs a support call.
Follow-up questions to expect
- "What does RAGAS faithfulness measure?" — It splits the answer into claims and asks a judge model whether each claim can be inferred from the retrieved context. Score = supported claims ÷ total claims.
- "Can you guarantee the model uses only the context?" — No. You can make it likely and detect when it does not. Say that plainly.
- "Why check citations in code if you have a judge?" — The code check is free and runs on every request; the judge costs an extra model call, so you run it on a sample or on high-risk answers.