Course Content
LLM Evaluation
6 sections · 50 lessons
How does textual entailment help verify correctness?
What you need to know
| Premise (policy text) | Hypothesis (claim in answer) | Label |
|---|---|---|
| Employees get 26 days of annual leave. | Employees get 26 days of leave a year. | Entailment |
| Employees get 26 days of annual leave. | Employees get 30 days of annual leave. | Contradiction |
| Employees get 26 days of annual leave. | Unused leave is paid out on exit. | Neutral |
An embedding model would score all three hypotheses as highly similar to the premise. NLI separates them.
Running an NLI model
1from transformers import pipeline23nli = pipeline("text-classification", model="microsoft/deberta-large-mnli")45premise = "Employees get 26 days of annual leave. Up to 8 days carry forward."6claim = "Employees can carry forward 10 days."7result = nli({"text": premise, "text_pair": claim})8print(result) # label is one of ENTAILMENT, NEUTRAL, CONTRADICTIONThe premise and hypothesis are passed as a sentence pair (text and text_pair), so the tokenizer inserts the model's own separator. Gluing them into one string with a hand-written [SEP] is a common bug: that token is not DeBERTa's separator.
Neutral vs contradiction
Treat them differently. A contradiction is a clear error. A neutral claim is unsupported — it may be true general knowledge ("leave requests go through the HR portal") or an invention. Many teams fail both in strict domains and only contradictions in relaxed ones.
Limits and workarounds
- Long premises — models trained on short sentence pairs degrade on multi-page context. Split the context into chunks and take the most favourable result across chunks for each claim.
- Multi-hop and numeric reasoning — "8 of 26 days carry forward" does not directly entail "18 days are lost if unused". Small NLI models miss this; LLM judges do better.
- Decompose first — a paragraph with four claims, one false, is often labelled as a whole. Check claim by claim.
- Domain shift — general NLI models trained on MNLI may be less accurate on legal or banking text; calibrate on your data.
A real-life example
The HR team compares two faithfulness checkers on 300 labelled claims: a DeBERTa NLI model (fast, runs on a CPU) and an LLM judge. On simple claims — one number, one condition — both reach about 0.9 accuracy. On claims needing a calculation or combining two clauses, the NLI model drops to 0.62 and the judge stays at 0.85.
They route by claim type: claims with a single number or date go to the NLI model; claims with "if", "unless" or arithmetic go to the judge. Cost falls by about 70% compared with judging everything, with accuracy close to the judge alone.
Follow-up questions to expect
- "Why not just use cosine similarity?" — Similarity measures topic closeness; "26 days" and "30 days" are nearly identical by cosine, but one contradicts the source.
- "How do you handle claims spread across two chunks?" — Chunk-level NLI misses them; use larger overlapping windows, retrieve the top chunks per claim and combine them, or use an LLM judge.
- "Is NLI better than an LLM judge?" — It is cheaper, faster and deterministic, and good for simple claims; LLM judges handle reasoning and long context better. Many systems use both.