Scenario-Based AI Engineering Questions

Course Content

Scenario-Based AI Engineering Questions

26 sections · 146 lessons

Your AI researcher cites sources confidently — but some citations don’t actually support the claims made. How do you build grounded citation systems that verify evidence before generation?


What you need to know

Why citations go wrong

A model writing an answer produces citations the same way it produces words: as likely-looking text. It can cite a real document that says something else, a document it never retrieved, or a real quote out of context or from an old version.

The verification pipeline

Python
def verify(claims, retrieved):    by_id = {d.id: d for d in retrieved}    results = []    for c in claims:                                   # c.text, c.doc_id, c.quote        doc = by_id.get(c.doc_id)        if doc is None:            results.append((c, "hallucinated_source"))       # not in retrieved set        elif normalise(c.quote) not in normalise(doc.text):            results.append((c, "quote_not_found"))        else:            results.append((c, nli(premise=c.quote, hypothesis=c.text)))  # entailed / contradicted / neutral    return results

The first two checks are simple lookups and cost almost nothing, yet they catch many bad citations. The entailment check compares one short claim with one short quote, so a small model can do it.

  1. Structured output — claims, each with a document id and an exact quote.
  2. Id check — the document must be in the retrieved set.
  3. Quote check — the quote must appear in that document.
  4. Entailment check — per claim.
  5. Act — entailed: keep. Not addressed: search again using the claim as the query; if still nothing, drop it or mark it unsupported. Contradicted: block and regenerate.
  6. Render honestly — the quote inline, unverified sentences visibly marked.

Put effective_date and document status in metadata and include them in the check, so a quote from a withdrawn guideline does not count as support.

Verification adds calls and latency; with a small checker model it is usually a modest share of the cost of a large model's answer. For research, medical and legal output it is clearly worth it. Some model APIs, such as Anthropic's citations feature, return the exact source passages used, which gives you the quote step directly.

A real-life example

Scenario, numbers made up. A pharma company's research assistant summarises clinical literature. A reviewer checks 200 answers and finds 11% of citations do not support their sentence; two cite papers that were never retrieved.

The team adds structured claims with quotes and the three checks. The id check catches 3% of citations, the quote check 4%, entailment another 5%. After re-retrieval, 70% of flagged claims find real support; the rest are dropped or marked "unsupported". In the next audit, unsupported citations fall to 1.5%, and answer latency rises by about 1.2 seconds, which reviewers accept.

Follow-up questions to expect

  • "What about a claim that combines two sources?" — Allow several quotes per claim and check entailment against them together; if it needs reasoning beyond the quotes, mark it as the assistant's inference.
  • "Can you check before generation, not after?" — Partly: extract supported facts from retrieved passages first, and have the model write only from those. You still verify the final text.
  • "How do you test the verifier?" — A labelled set of claim–quote pairs, including tricky ones like negations and number changes.