Course Content
LLM Evaluation
6 sections · 50 lessons
What is the difference between local and global factual consistency?
What you need to know
Local consistency
- Unit: one claim or sentence
- Question: is this supported by the source?
- Tools: claim decomposition plus NLI or judge
- Localises the error exactly
Global consistency
- Unit: the whole output
- Question: does it represent the source fairly?
- Tools: document-level judge, coverage checklists
- Catches omissions, contradictions, wrong emphasis
Failures only a global check catches
- Omission that changes meaning — the source says "refund approved, subject to merchant confirmation"; the summary says "refund approved". Every word is supported; the conclusion is wrong.
- Self-contradiction — sentence 2 says the complaint was resolved, sentence 5 says it is pending. Each is supported by a different part of a long thread, but together they confuse.
- Entity mixing — two customers in one email thread; the summary gives one customer's account number with the other's complaint.
- Wrong emphasis — a minor issue presented as the main one.
Local checks miss all of these because each sentence is verified against the source independently, never against the other sentences or against what is missing.
How to check globally
- A document-level judge with specific questions: "Does the summary contradict itself?", "Does it omit any condition that changes the outcome?", "Are all facts attributed to the right person?".
- A coverage checklist — list the source's key points (by hand for the golden set, or with a model) and check each appears.
- Entity-level checks — every account number or name appears with the right attributes.
A real-life example
The bank summarises long complaint email threads for supervisors. On 100 threads, claim-level faithfulness is 0.97 — almost every sentence is supported. Yet supervisors report wrong decisions.
A global check finds the cause. In 11 of 100 summaries, a condition that changes the outcome was dropped — "refund approved, pending chargeback result" became "refund approved". In 4, two customers in a forwarded thread were merged. The team adds a checklist of required elements (status, amount, pending conditions, customer ID) and a judge question for contradictions. The dashboard now shows local faithfulness and a "key-condition coverage" rate side by side; coverage becomes the metric supervisors care about.
Follow-up questions to expect
- "Can one LLM judge do both checks?" — It can, but separate questions work better: claim-by-claim for local, and a few specific global questions, each calibrated against human labels.
- "How do you evaluate omissions without a reference summary?" — Extract required points from the source (by rule, by a model, or by hand for a golden set) and check coverage.
- "Does longer context make global consistency harder?" — Yes; judges and generators both miss details in long sources, so chunked checks plus a coverage list are more reliable.