Course Content
AI Safety & Guardrails
5 sections · 50 lessons
What is grounding, and how does it reduce hallucinations?
What you need to know
The four parts of a grounded pipeline
- Retrieve — relevant, permission-filtered, versioned evidence. The assistant sees only documents this user may see.
- Instruct — answer only from the evidence, and cite the passage for each claim.
- Allow abstention — if the evidence does not answer the question, say so or ask a follow-up.
- Verify — check each claim is supported by its cited passage before returning it.
Skipping any step is where teams go wrong. Without step 3, a model given poor retrieval fills the gap from memory and sounds exactly the same.
Verification methods
- Quote checking: require verbatim quotes and check they exist in the source. Cheap and deterministic.
- Entailment (NLI): a small model scores whether the passage entails the sentence.
- LLM judge: a model grades each claim as supported, contradicted or not found.
- Managed checks: some platforms offer this, for example the contextual grounding check in AWS Bedrock Guardrails and groundedness detection in Azure AI Content Safety.
1import re23def norm(s):4 return re.sub(r"\s+", " ", s).strip().lower()56def unsupported_quotes(answer, sources):7 corpus = [norm(s) for s in sources]8 quotes = re.findall(r'"([^"]{12,})"', answer)9 return [q for q in quotes if not any(norm(q) in c for c in corpus)]1011sources = ["Foreclosure of a floating-rate home loan attracts no "12 "prepayment charge for individual borrowers."]13answer = ('Per policy, "foreclosure of a floating-rate home loan attracts no '14 'prepayment charge". Fixed-rate loans carry "a 2% charge on the '15 'outstanding principal".')16print(unsupported_quotes(answer, sources))17# ['a 2% charge on the outstanding principal']The function finds every quoted span in the answer and checks it appears in a source. The second quote was invented, so the answer is blocked or regenerated. Quote checking only proves the quote exists; an NLI or judge step is still needed to check the sentence around it.
Failure modes to name
- Retrieval miss: the right document is not retrieved; without abstention the model guesses.
- Stale or conflicting sources: grounding faithfully repeats an outdated policy.
- Citation laundering: a citation is attached to a sentence the source does not support. Measure citation precision, not just whether citations exist.
- Injected sources: a retrieved document contains instructions. Grounding in untrusted text needs injection defences.
A real-life example
A healthcare symptom-checker answers questions from a library of 1,200 clinician-approved articles. Before grounding, it answered "Can I take ibuprofen with dengue?" with "Yes, in normal doses" — dangerous, because NSAIDs are avoided in dengue due to bleeding risk.
After grounding, the retrieved article says to use paracetamol and avoid ibuprofen and aspirin, and the answer quotes it. Verification catches a different problem in testing: for 4% of answers the model added a dose that no article contained. Those answers are now blocked and regenerated with a stricter instruction; if they fail again, the bot shows the article link and suggests seeing a doctor.
Follow-up questions to expect
- "How is grounding different from fine-tuning?" — Fine-tuning changes the weights and cannot be cited or updated quickly. Grounding supplies evidence per request; update a document and the next answer changes.
- "What if the documents disagree?" — Prefer the newest version by metadata, show both with dates, or abstain. Resolving conflicts is a data-management job before it is a model job.
- "How do you measure it?" — Faithfulness score, citation precision, and abstention rate on questions the corpus cannot answer.