Course Content
Applied AI Engineering: From Prompt to Production
9 sections · 29 lessons
Grounding and citations
With recall@5 at 0.91, the right policy text is in the prompt for nine questions out of ten. The team assumed the answers would now be right nine times out of ten too. Then a reviewer read 200 answers sentence by sentence and found that 14% of sentences contained a claim that no cited source supported.
Most were small. "You must inform your manager two weeks in advance" was added to a leave answer, though the policy says nothing about notice. A sentence cited source [2] for a fact that was in source [3]. A travel answer said "economy class is standard for all flights", which is common practice but not what Harbourline's policy says for flights over eight hours. Each sentence sounded reasonable. That is precisely the danger: the unsupported sentences sounded exactly as confident as the supported ones.
Grounding means every claim in the answer comes from a source the system can point to. Retrieval makes grounding possible. It does not make it happen. This lesson makes it happen, and then checks it in code.
Grounded answers start in the prompt
The version 1 prompt already said "use only the text inside the sources" and "every factual sentence must cite a source". The reviewer's findings led to three sharper rules.
- Cite at the end of each sentence that states a fact, not once at the end of the answer.
- If a source does not say something, do not say it, even if it is usually true at other companies.
- If sources seem to conflict, say so and cite both, rather than choosing silently.
These changes, plus the dynamic examples from Section 2, brought unsupported sentences from 14% to 6%. That is a large improvement, and still far too high for a system people use to make decisions about their leave and money. Prompt rules reduce a failure; they do not remove it. The rest has to be caught in code.
Check citations in code
PolicyPal checks each answer in three steps. The first two are cheap string checks. The third uses a small natural language inference (NLI) model, which reads a premise and a hypothesis and says whether the premise entails, contradicts or is neutral towards the hypothesis. Here the premise is the cited chunk and the hypothesis is the answer sentence.
1# policypal/grounding.py2import re34from sentence_transformers import CrossEncoder56nli = CrossEncoder("cross-encoder/nli-deberta-v3-small")7LABELS = ["contradiction", "entailment", "neutral"] # this model's label order8CITE = re.compile(r"\[(\d+)\]")910def sentences(text: str) -> list[str]:11 return [s.strip() for s in re.split(r"(?<=[.!?])\s+", text) if s.strip()]1213def check(answer: str, sources: list[dict]) -> list[dict]:14 """Returns one finding per sentence that is not clearly supported."""15 findings, pairs, where = [], [], []16 for sent in sentences(answer):17 ids = [int(n) for n in CITE.findall(sent)]18 if not ids:19 findings.append({"sentence": sent, "problem": "no citation"})20 continue21 claim = CITE.sub("", sent).strip()22 for n in ids:23 if 1 <= n <= len(sources):24 pairs.append((sources[n - 1]["text"], claim))25 where.append(sent)26 if pairs:27 labels = [LABELS[i] for i in nli.predict(pairs).argmax(axis=1)]28 supported = {s for s, lab in zip(where, labels) if lab == "entailment"}29 for sent in dict.fromkeys(where):30 if sent not in supported:31 findings.append({"sentence": sent, "problem": "not supported by its citation"})32 return findingsThe NLI model runs on the CPU and checks a typical three-sentence answer in about 90 milliseconds. A sentence counts as supported if any of its cited sources entails it. Sentences without a citation are flagged immediately. A sentence like "Please contact HR for more help" would be flagged too, so in practice PolicyPal skips sentences that match a short list of known guidance phrases.
Why a small NLI model and not a second LLM call? It is faster, it costs nothing per call, and it is deterministic, so the same answer always gets the same verdict. Its weakness is paraphrase: it is strict, and it flags about 4% of sentences that a human would call supported, mostly where the answer summarised a long rule. An LLM checker handles paraphrase better, at a cost of 300 to 800 milliseconds and a fraction of a cent per answer. PolicyPal uses NLI on every answer, and Section 6 uses an LLM judge offline for deeper grading.
What to do with an unsupported sentence
A flagged sentence needs a policy, decided in advance.
- Drop it if the rest still answers — if the unsupported sentence is extra detail and the remaining sentences still answer the question, remove it and show the rest.
- Regenerate once with the finding — if the unsupported sentence is the main claim, send the finding back ("sentence 2 is not supported by source 3") and ask again.
- Hand over if it still fails — set status to
needs_humanand route the question to HR, with the sources attached.
The first step handles most cases, because unsupported sentences are usually additions, like the two-week notice rule. Regenerating costs another model call, so it is reserved for the core claim. Because the NLI model is strict, the regenerate-then-hand-over path is better than simply blocking: a false alarm usually passes on the second try, when the model quotes the source more closely.
Stale and conflicting sources
Some wrong answers were perfectly grounded in the wrong source. A superseded travel policy said one thing; the current one said another. The answer faithfully cited the old one. No citation check catches this, because the citation is correct.
The fix belongs in the data and retrieval layers, not the prompt.
- Index only current versions, as set up in the chunking lesson, and remove a superseded version from the index on the day it stops applying.
- Keep both versions during a transition, with
effective_fromon each chunk, and show that date in the source header, so the model can say "until 31 March ... from 1 April ...". - Prefer the specific over the general. When an India policy and a Global standard both match, the retrieval step keeps the country-specific chunk above the global one, because Harbourline's rule is that local policy overrides global.
Encoding these rules in retrieval means the model rarely sees a conflict it has to resolve. Asking a model to "use the newest version" works most of the time; making sure only the right version is present works all the time.
Show the sources to the user
Finally, grounding is also a user interface. Each citation in PolicyPal's answer links to the exact PDF and page. Employees quickly learned to click them for anything that mattered, such as money or deadlines. That habit is a safety net no check can replace, and it builds the right kind of trust: trust in a system that shows its evidence, not in one that merely sounds sure.
Check your understanding
0 of 3 answered
1.An answer cites a superseded travel policy correctly, and the NLI check marks every sentence as supported. What is the right fix?
2.Why does PolicyPal use a small NLI model on every answer instead of asking a second LLM to check support?
3.The NLI check flags one sentence out of four: "You must give your manager two weeks' notice." The other three sentences fully answer the question. What should happen?