Live Coding Interview Prep

Course Content

Live Coding Interview Prep

7 sections · 50 lessons

Add citation tracking to a RAG pipeline.


Markers parsed from one answer, two chunks in the prompt[1][2][1][7]0123pol-12, page 3faq-4seenalready: skipno chunk7: invalid
Only markers that resolve to a real chunk become links; the out-of-range one is a hallucinated citation to log.

What you need to know

A citation is useful only if it is verifiable: a user clicks [2] and sees the page it came from. That needs three pieces:

  • Stable ids. Every chunk carries its id, source file and page in metadata, not only in its text.
  • A numbered prompt. The model sees [1] …, [2] … and is told to cite them. Numbers are short and easy for the model to copy correctly; long ids are not.
  • Parse and resolve. After generation, find every [n], check that n is in range, de-duplicate, and map it to the chunk.

Two quality checks go on top:

  • Uncited claims — a factual sentence with no marker is unsupported. Flag it or block the answer.
  • Citation verification — the model can cite the wrong chunk. A cheap NLI model or a small LLM call ("does chunk 2 support this sentence?") catches it.
Python
import refrom collections.abc import Callablefrom dataclasses import dataclass@dataclassclass Chunk:    id: str    text: str    source: str    page: int | None = NoneCITE_RE = re.compile(r"\[(\d+)\]")def answer_with_citations(query: str, chunks: list[Chunk], llm_fn: Callable[[str], str]) -> dict:    """Generate an answer and resolve its [n] markers to real sources."""    if not chunks:        return {"answer": "I have no sources for that.", "citations": [], "invalid": []}    context = "\n\n".join(f"[{n}] {c.text}" for n, c in enumerate(chunks, 1))    text = llm_fn(        "Answer using only the numbered context. Put the source number as [n] after "        "each factual claim. If the context is insufficient, say so.\n\n"        f"Context:\n{context}\n\nQuestion: {query}\nAnswer:"    )    citations, seen, invalid = [], set(), set()    for m in CITE_RE.finditer(text):        n = int(m.group(1))        if not 1 <= n <= len(chunks):            invalid.add(n)                       # the model cited a source that doesn't exist        elif n not in seen:            seen.add(n)            c = chunks[n - 1]            citations.append({"marker": n, "chunk_id": c.id, "source": c.source, "page": c.page})    return {"answer": text, "citations": citations, "invalid": sorted(invalid)}

The tricky parts:

  • 1 <= n <= len(chunks) — markers are 1-based in the prompt and the list is 0-based, hence chunks[n - 1]. [0] is invalid too.
  • seen keeps first-use order: the first source mentioned becomes citation 1 in the UI.
  • finditer handles grouped markers like [1][3] with no extra code.

Complexity: building the prompt is O(total chunk length); parsing is O(answer length); resolving is O(1) per marker. Space O(number of markers).

A real-life example

Python
chunks = [Chunk("pol-12", "Refunds take 5 working days.", "refund-policy.pdf", 3),          Chunk("faq-4", "Refunds need the order id.", "faq.md")]fake_llm = lambda _: "A refund takes 5 working days [1] and needs the order id [2][1]. Call us [7]."out = answer_with_citations("How do refunds work?", chunks, fake_llm)print(out["citations"])# [{'marker': 1, 'chunk_id': 'pol-12', 'source': 'refund-policy.pdf', 'page': 3},#  {'marker': 2, 'chunk_id': 'faq-4', 'source': 'faq.md', 'page': None}]print(out["invalid"])   # [7]

The regex finds four markers in order: [1], [2], [1], [7].

markerin range 1–2?seen before?result
1yesnocitation → pol-12, page 3
2yesnocitation → faq-4
1yesyesskipped (duplicate)
7no–added to invalid

The UI renders [1] and [2] as links and shows [7] as a warning or strips the sentence.

A bank's policy assistant uses this so every answer about loan charges links to the exact page of the rate card — auditors check the page, not the chatbot.

Follow-up questions to expect

  • "The model cites the right number but the chunk does not support the claim — how do you catch that?" — Verify each cited sentence against its chunk with an NLI model or a small LLM judge, and drop or flag unsupported sentences.
  • "What about providers with built-in citations?" — Some APIs, including Anthropic's, can return citations with exact character or page locations when you pass documents with citations enabled. That removes the parsing step; the resolution and display logic stay the same.
  • "What if you reorder chunks after numbering?" — Renumber after any reordering, or every marker points at the wrong source.