Applied AI Engineering: From Prompt to Production

Course Content

Applied AI Engineering: From Prompt to Production

9 sections · 29 lessons

Grounding and citations


With recall@5 at 0.91, the right policy text is in the prompt for nine questions out of ten. The team assumed the answers would now be right nine times out of ten too. Then a reviewer read 200 answers sentence by sentence and found that 14% of sentences contained a claim that no cited source supported.

Most were small. "You must inform your manager two weeks in advance" was added to a leave answer, though the policy says nothing about notice. A sentence cited source [2] for a fact that was in source [3]. A travel answer said "economy class is standard for all flights", which is common practice but not what Harbourline's policy says for flights over eight hours. Each sentence sounded reasonable. That is precisely the danger: the unsupported sentences sounded exactly as confident as the supported ones.

Grounding means every claim in the answer comes from a source the system can point to. Retrieval makes grounding possible. It does not make it happen. This lesson makes it happen, and then checks it in code.

Making every sentence traceablePrompt: cite each factual sentenceCode: every citation id existsNLI: cited text entails the sentenceDrop it, regenerate once, or hand over
Prompt rules took unsupported sentences from 14% to 6%; only a check in code brought them to 3%.

Grounded answers start in the prompt

The version 1 prompt already said "use only the text inside the sources" and "every factual sentence must cite a source". The reviewer's findings led to three sharper rules.

  • Cite at the end of each sentence that states a fact, not once at the end of the answer.
  • If a source does not say something, do not say it, even if it is usually true at other companies.
  • If sources seem to conflict, say so and cite both, rather than choosing silently.

These changes, plus the dynamic examples from Section 2, brought unsupported sentences from 14% to 6%. That is a large improvement, and still far too high for a system people use to make decisions about their leave and money. Prompt rules reduce a failure; they do not remove it. The rest has to be caught in code.

Check citations in code

PolicyPal checks each answer in three steps. The first two are cheap string checks. The third uses a small natural language inference (NLI) model, which reads a premise and a hypothesis and says whether the premise entails, contradicts or is neutral towards the hypothesis. Here the premise is the cited chunk and the hypothesis is the answer sentence.

Python
# policypal/grounding.pyimport refrom sentence_transformers import CrossEncodernli = CrossEncoder("cross-encoder/nli-deberta-v3-small")LABELS = ["contradiction", "entailment", "neutral"]      # this model's label orderCITE = re.compile(r"\[(\d+)\]")def sentences(text: str) -> list[str]:    return [s.strip() for s in re.split(r"(?<=[.!?])\s+", text) if s.strip()]def check(answer: str, sources: list[dict]) -> list[dict]:    """Returns one finding per sentence that is not clearly supported."""    findings, pairs, where = [], [], []    for sent in sentences(answer):        ids = [int(n) for n in CITE.findall(sent)]        if not ids:            findings.append({"sentence": sent, "problem": "no citation"})            continue        claim = CITE.sub("", sent).strip()        for n in ids:            if 1 <= n <= len(sources):                pairs.append((sources[n - 1]["text"], claim))                where.append(sent)    if pairs:        labels = [LABELS[i] for i in nli.predict(pairs).argmax(axis=1)]        supported = {s for s, lab in zip(where, labels) if lab == "entailment"}        for sent in dict.fromkeys(where):            if sent not in supported:                findings.append({"sentence": sent, "problem": "not supported by its citation"})    return findings

The NLI model runs on the CPU and checks a typical three-sentence answer in about 90 milliseconds. A sentence counts as supported if any of its cited sources entails it. Sentences without a citation are flagged immediately. A sentence like "Please contact HR for more help" would be flagged too, so in practice PolicyPal skips sentences that match a short list of known guidance phrases.

Why a small NLI model and not a second LLM call? It is faster, it costs nothing per call, and it is deterministic, so the same answer always gets the same verdict. Its weakness is paraphrase: it is strict, and it flags about 4% of sentences that a human would call supported, mostly where the answer summarised a long rule. An LLM checker handles paraphrase better, at a cost of 300 to 800 milliseconds and a fraction of a cent per answer. PolicyPal uses NLI on every answer, and Section 6 uses an LLM judge offline for deeper grading.

What to do with an unsupported sentence

A flagged sentence needs a policy, decided in advance.

  1. Drop it if the rest still answers — if the unsupported sentence is extra detail and the remaining sentences still answer the question, remove it and show the rest.
  2. Regenerate once with the finding — if the unsupported sentence is the main claim, send the finding back ("sentence 2 is not supported by source 3") and ask again.
  3. Hand over if it still fails — set status to needs_human and route the question to HR, with the sources attached.

The first step handles most cases, because unsupported sentences are usually additions, like the two-week notice rule. Regenerating costs another model call, so it is reserved for the core claim. Because the NLI model is strict, the regenerate-then-hand-over path is better than simply blocking: a false alarm usually passes on the second try, when the model quotes the source more closely.

Stale and conflicting sources

Some wrong answers were perfectly grounded in the wrong source. A superseded travel policy said one thing; the current one said another. The answer faithfully cited the old one. No citation check catches this, because the citation is correct.

The fix belongs in the data and retrieval layers, not the prompt.

  • Index only current versions, as set up in the chunking lesson, and remove a superseded version from the index on the day it stops applying.
  • Keep both versions during a transition, with effective_from on each chunk, and show that date in the source header, so the model can say "until 31 March ... from 1 April ...".
  • Prefer the specific over the general. When an India policy and a Global standard both match, the retrieval step keeps the country-specific chunk above the global one, because Harbourline's rule is that local policy overrides global.

Encoding these rules in retrieval means the model rarely sees a conflict it has to resolve. Asking a model to "use the newest version" works most of the time; making sure only the right version is present works all the time.

Show the sources to the user

Finally, grounding is also a user interface. Each citation in PolicyPal's answer links to the exact PDF and page. Employees quickly learned to click them for anything that mattered, such as money or deadlines. That habit is a safety net no check can replace, and it builds the right kind of trust: trust in a system that shows its evidence, not in one that merely sounds sure.

Check your understanding

0 of 3 answered

1.An answer cites a superseded travel policy correctly, and the NLI check marks every sentence as supported. What is the right fix?

2.Why does PolicyPal use a small NLI model on every answer instead of asking a second LLM to check support?

3.The NLI check flags one sentence out of four: "You must give your manager two weeks' notice." The other three sentences fully answer the question. What should happen?