Course Content
LLM Evaluation
6 sections · 50 lessons
What is the difference between token-level and document-level grounding?
What you need to know
The levels
| Level | What is linked | Example | Catches |
|---|---|---|---|
| Document | Whole answer to a document | "Source: Leave Policy 2026.pdf" | Answers from the wrong document |
| Passage or sentence | Each sentence to a passage | "Up to 8 days carry forward [Leave Policy, section 4.2]" | Wrong or invented facts within a sentence |
| Span or token | Phrases to exact character ranges | Highlighted quote in the source | Precise misquotes |
The classic failure document-level misses
The answer cites the correct policy, but says "10 days carry forward" while the policy says 8. A document-level check passes: the right document was used. A sentence-level check with entailment fails that sentence.
How span-level grounding is produced
- Prompting for citations — ask for a citation after each sentence, pointing to numbered chunks. Simple, but citations can be wrong.
- Provider citation features — some model APIs can return citations that point to exact passages of documents you supply, which is more reliable than free-text citation markers.
- Extractive or quote-first generation — the model first copies exact quotes, then answers only from them; code checks each quote exists in the source.
- Attribution analysis — research methods using attention or gradients; rarely used in production.
Evaluating citations
Measure two things (terms used in attribution research):
- Citation precision — of the citations given, how many actually support their sentence?
- Citation recall — of the sentences that need support, how many have a supporting citation?
Both are checked with NLI or a judge: does the cited passage entail the sentence?
A real-life example
The HR assistant used document-level citations: every answer ended with the policy file name. An audit of 200 answers found the right document cited 96% of the time, yet 9% of answers had a wrong number or condition.
The team switched to sentence-level citations with a quote-first step: the model quotes the lines it relies on, the code checks each quote exists in the retrieved chunks, and an NLI check confirms each sentence is entailed by its quote. Citation precision is 0.94 and recall 0.91 on the audit set, and wrong-number answers fall to 2%. Employees can now click a sentence and see the highlighted policy line — which the compliance team had asked for.
Follow-up questions to expect
- "Why not always use span-level grounding?" — It adds prompt complexity, latency and more ways to fail; for low-risk surfaces, document-level citations may be enough.
- "Can the model's citations be trusted?" — Not without checking; models cite plausible passages that don't support the sentence. Verify every citation with entailment or quote matching.
- "How does chunking affect grounding?" — Very small chunks lose context; very large chunks make citations vague. Sentence-level citations work best when chunks keep sections intact.