Evaluating and Testing GenAI Models

Retrieval-Based Consistency Checks


A retrieval-augmented support bot is given a policy document and asked whether a customer can get a refund after 45 days. The document says: "Refunds are available within 30 days of purchase. Enterprise customers may request an extension of up to 60 days, subject to approval."

The bot answers: "Yes, refunds are available for up to 60 days, so a 45-day-old purchase is eligible."

Every fact in that sentence appears in the source. The number 60 is there. The word "refunds" is there. A similarity metric scores this answer highly, because it overlaps heavily with the document. A human reads it and immediately sees the failure: the 60-day extension is conditional on being an Enterprise customer and on approval, and the bot dropped both conditions. The answer is unsupported despite being made entirely of supported pieces.

This is the gap that retrieval-based consistency checking exists to close. The question is not "does this text look like the source" — it is "does the source entail this specific claim". Those are different relations, and only the second one is what grounding actually promises.

Verifying the refund answer, claim by claimResponseSplit intoatomic claimsRetrieveevidenceper claimNLI: entails,contradicts, neutralVerdict plusthe citationRetrieve with the claim, not the question — the retriever caps every metric downstream.
The relation you need is entailment, not similarity: the policy passage is highly similar to the false 45-day answer and still refutes it.

The core relation: entailment, not similarity

Natural language inference (NLI) is the task of deciding, given a premise and a hypothesis, which of three relations holds:

LabelMeaningExample premiseExample hypothesis
EntailmentIf the premise is true, the hypothesis must be trueRefunds are available within 30 days of purchase.A purchase made 20 days ago can be refunded.
ContradictionIf the premise is true, the hypothesis must be falseRefunds are available within 30 days of purchase.Refunds are available for any purchase regardless of age.
NeutralThe premise neither guarantees nor rules out the hypothesisRefunds are available within 30 days of purchase.Most customers request refunds within a week.

Apply it to the opening failure. Premise: the policy paragraph. Hypothesis: "A 45-day-old purchase is eligible for a refund." A decent NLI model returns neutral or contradiction, because the premise only licenses that conclusion under conditions the hypothesis does not state. Cosine similarity between the two texts, meanwhile, is around 0.8. The two methods disagree, and the entailment answer is the correct one.

Similarity asks whether two texts are about the same thing. Entailment asks whether one guarantees the other. Grounding is a promise about the second, so measuring the first is measuring the wrong relation.

The models used are typically RoBERTa or DeBERTa fine-tuned on MNLI, or a purpose-trained faithfulness model. They cost a few milliseconds per pair on a GPU, which is what makes checking every claim in every response affordable.

The pipeline

Text
Response   |[1] Decompose into atomic claims   |     "Refunds up to 60 days" + "45-day purchase is eligible"   |     (a compound sentence hides which half is unsupported)   v[2] Retrieve evidence per claim, k passages   |     query = the claim itself, not the original user question   v[3] Score each (passage, claim) pair with NLI   |     keep max entailment and max contradiction across passages   v[4] Decide: SUPPORTED / CONTRADICTED / NOT_ENOUGH_INFO   |[5] Aggregate to response-level and corpus-level rates

Three of those five steps are where the errors come from.

Step 1: decomposition is not optional

Run NLI on a whole paragraph and you get one label for a text containing six claims, five supported and one fabricated. The label will usually be "neutral", which tells you nothing. Atomic decomposition is what turns a vague verdict into a specific, fixable finding — and it also fixes the partial-truth problem, where a claim is 80% right.

Decompose so that each unit contains one predicate and its arguments, with any conditions attached:

Text
BAD  (one claim):  "Refunds are available for up to 60 days, so a 45-day-old                    purchase is eligible."GOOD (three claims):  c1  Refunds have a standard window of 60 days.              -> CONTRADICTED (it is 30)  c2  A 60-day window requires Enterprise status and approval. -> not asserted; omission  c3  A 45-day-old purchase is eligible for a refund.          -> NOT ENTAILED

Step 2: retrieve with the claim, not the question

A common bug: the pipeline retrieves once using the user's question and reuses those passages to verify every claim. But the claims a model produces often concern details the question did not mention. Re-retrieving per claim, using the claim text as the query, typically raises evidence recall by 10 to 20 points on any realistic corpus. It costs one extra search per claim, which is cheap relative to the generation that produced it.

Step 3: the retriever bounds everything downstream

This is the most under-appreciated fact in the whole area. Your verification recall cannot exceed your retrieval recall. If the retriever finds the relevant passage 82% of the time and your NLI model is 91% accurate given the right passage, the end-to-end ceiling is:

recallend-to-end=0.82×0.91=0.746\text{recall}_{\text{end-to-end}} = 0.82 \times 0.91 = 0.746

You can spend a month upgrading the NLI model from 91% to 95% and gain 0.82×0.04=3.30.82 \times 0.04 = 3.3 points. Raising retrieval recall from 82% to 92% gains 0.10×0.91=9.10.10 \times 0.91 = 9.1 points, for less work. Measure recall@k on a labelled evidence set before you touch anything else, because it tells you which half of the system to invest in.

Making retrieval good enough

Sparse, dense, and why you want both

Sparse (BM25)Dense (embeddings)
Matches onExact terms, weighted by raritySemantic proximity in vector space
Excellent forProduct codes, error strings, names, numbers, rare jargonParaphrase, synonymy, conceptual queries
Fails on"Cancel my subscription" vs "terminate the plan"Exact identifiers: ORD-4471 embeds near ORD-4472
Needs training dataNoYes, if you fine-tune the encoder
Index costLowHigher; re-embedding on model change

Fact-checking queries are unusually adversarial for dense retrieval, because the highest-risk claims are exactly the ones full of numbers and identifiers — the case where sparse retrieval is much stronger. So combine them with reciprocal rank fusion, which needs no score calibration:

RRF(d)=∑r∈retrievers1k+rankr(d)k=60\text{RRF}(d) = \sum_{r \in \text{retrievers}} \frac{1}{k + \text{rank}_r(d)} \qquad k = 60

Worked example with two documents:

DocumentBM25 rankDense rankRRF score
D1312163+172=0.01587+0.01389=0.02976\frac{1}{63} + \frac{1}{72} = 0.01587 + 0.01389 = 0.02976
D24011100+161=0.01000+0.01639=0.02639\frac{1}{100} + \frac{1}{61} = 0.01000 + 0.01639 = 0.02639

D1 wins despite never being either retriever's top hit, because it is consistently good. That is the behaviour you want when the two retrievers have complementary blind spots. The constant k=60k = 60 damps the influence of the top rank so a single retriever's confident mistake cannot dominate.

Python
from collections import defaultdictdef rrf(ranked_lists, k: int = 60, top_n: int = 10):    """ranked_lists: list of lists of doc ids, each already in rank order."""    scores = defaultdict(float)    for lst in ranked_lists:        for rank, doc_id in enumerate(lst, start=1):            scores[doc_id] += 1.0 / (k + rank)    return sorted(scores.items(), key=lambda kv: -kv[1])[:top_n]

Chunking decides what can be verified

If a claim requires two sentences that your chunker split apart, no retriever can supply entailing evidence and the claim will be labelled NEI forever. Practical rules: chunk on semantic boundaries rather than fixed token counts; overlap chunks by 15–20%; and keep a parent-document reference so the verifier can widen its window when the top chunk is nearly sufficient.

Turning NLI scores into verdicts

The NLI model returns a distribution over three labels. Turning that into a verdict requires thresholds, and the thresholds are a genuine engineering decision, not a default to accept.

Python
from dataclasses import dataclass@dataclassclass Check:    claim: str    verdict: str    entail: float    contradict: float    evidence_id: str | Nonedef check_claim(claim, retriever, nli, k=5, tau_e=0.75, tau_c=0.70) -> Check:    passages = retriever.search(claim, k=k)    if not passages:        return Check(claim, "NOT_ENOUGH_INFO", 0.0, 0.0, None)    best_e, best_c, best_id = 0.0, 0.0, None    for p in passages:        s = nli(premise=p.text, hypothesis=claim)        if s["entailment"] > best_e:            best_e, best_id = s["entailment"], p.id        best_c = max(best_c, s["contradiction"])    if best_c >= tau_c and best_c > best_e:        return Check(claim, "CONTRADICTED", best_e, best_c, best_id)    if best_e >= tau_e:        return Check(claim, "SUPPORTED", best_e, best_c, best_id)    return Check(claim, "NOT_ENOUGH_INFO", best_e, best_c, best_id)

Note that contradiction is checked first. A claim with entailment 0.78 and contradiction 0.85 against different passages means your corpus disagrees with itself, or the claim is subtly wrong in a way one passage catches. Treating it as supported because the entailment threshold was met would hide the more important signal.

Choosing the threshold with data, not intuition

Label 200 claims by hand, then sweep:

τe\tau_ePrecision of SUPPORTEDRecall of SUPPORTEDNEI rateReading
0.500.790.966%Rubber-stamps almost everything
0.650.870.9111%Reasonable default
0.750.930.8319%Good when a false pass is costly
0.850.970.6832%Nearly a third unresolved — review queue floods
0.950.990.4157%Unusable

Pick the row by asking what happens to a false SUPPORTED in your product. If it goes straight to a customer as a policy statement, take 0.85 and staff the review queue. If it feeds an internal quality dashboard, take 0.65 and keep the noise floor low.

The metrics, computed on a real batch

One batch: 40 responses, decomposed into 168 claims. Verdicts: 141 supported, 9 contradicted, 18 not enough info.

Faithfulness=supportedtotal=141168=0.839\text{Faithfulness} = \frac{\text{supported}}{\text{total}} = \frac{141}{168} = 0.839

Contradiction rate=9168=0.054Unverifiable rate=18168=0.107\text{Contradiction rate} = \frac{9}{168} = 0.054 \qquad \text{Unverifiable rate} = \frac{18}{168} = 0.107

Report all three. They point at different bugs: contradiction rate is a generator problem, unverifiable rate is usually a retrieval or corpus-coverage problem, and only their sum shows up in faithfulness.

Response-level metrics matter too, because a user experiences responses, not claims. If 31 of the 40 responses have every claim supported:

Fully-grounded response rate=3140=0.775\text{Fully-grounded response rate} = \frac{31}{40} = 0.775

The gap between 83.9% of claims and 77.5% of responses is informative: unsupported claims are somewhat concentrated, but not entirely — they are spread across 9 responses rather than clustered in 2.

The interval, computed correctly

The 168 claims come from 40 responses, so they are clustered and the naive standard error understates the noise. With average cluster size m=4.2m = 4.2 and intra-cluster correlation ρ=0.25\rho = 0.25:

DEFF=1+(m−1)ρ=1+3.2×0.25=1.80neff=1681.80=93\text{DEFF} = 1 + (m-1)\rho = 1 + 3.2 \times 0.25 = 1.80 \qquad n_{\text{eff}} = \frac{168}{1.80} = 93

SE=0.839×0.16193=0.001453=0.0381SE = \sqrt{\frac{0.839 \times 0.161}{93}} = \sqrt{0.001453} = 0.0381

So faithfulness is 83.9%±7.5%83.9\% \pm 7.5\% at 95% confidence, roughly [76.4%, 91.4%][76.4\%,\ 91.4\%]. Anyone comparing this to a 87.5% run from last week and declaring a regression is reading noise. To resolve a 4-point difference in faithfulness at this variance you need roughly neff≈1,300n_{\text{eff}} \approx 1{,}300, meaning about 2,300 raw claims, meaning about 550 responses per arm.

Every consistency metric you compute is a proportion over clustered units. If your dashboard shows faithfulness to one decimal place with no interval, it is showing precision it does not have.

Self-consistency: checking without any source

Sometimes there is no document to check against. A useful proxy — the idea behind SelfCheckGPT — is that a model producing a fact it actually encoded will reproduce it across independent samples, while a fabrication will vary.

Sample the same prompt kk times at non-zero temperature, then measure agreement between the samples using the same NLI machinery, one sample as premise and the target claim as hypothesis.

Worked example with k=5k = 5 samples, checking one claim against the other four, plus all pairwise comparisons — 10 pairs in total. If 7 pairs are mutually entailing:

Consistency=710=0.70\text{Consistency} = \frac{7}{10} = 0.70

Claims below about 0.5 consistency are strong hallucination candidates. Two hard limits, though:

  • Cost. kk samples means kk times the generation spend. Practical only on a sample of traffic or on flagged responses.
  • Stable falsehoods pass. A misconception the model has firmly encoded reproduces perfectly across samples and scores 1.0 consistency. Self-consistency detects uncertainty, not error, and those overlap only partially.

Multi-hop claims, where the arithmetic turns against you

Some claims cannot be verified by any single passage. "The company's current CTO previously worked at the firm that acquired our largest supplier" requires three separate retrievals chained together. Two things go wrong.

Retrieval cannot find the second hop from the original query. The passage naming the acquirer contains none of the terms in the claim about the CTO. You need iterative retrieval: verify hop one, extract the entity it resolves, use that entity as the query for hop two.

Errors compound multiplicatively. If each hop is verified correctly with probability 0.85:

HopsProbability the whole chain is right
10.850.85
20.852=0.7230.85^2 = 0.723
30.853=0.6140.85^3 = 0.614
40.854=0.5220.85^4 = 0.522

At three hops your verifier is right about six times in ten — barely better than a coin flip on a binary decision. This is not fixable by improving the NLI model; going from 0.85 to 0.92 per hop still only reaches 0.923=0.7790.92^3 = 0.779. The practical response is to report multi-hop claims as a separate category with their own (much wider) uncertainty, rather than pooling them into a single faithfulness number where they silently degrade it.

Python
def verify_chain(claim_steps, retriever, nli):    """Verify a multi-hop claim step by step, carrying resolved entities forward."""    context, trace = {}, []    for step in claim_steps:        query = step["query_template"].format(**context)        result = check_claim(query, retriever, nli)        trace.append(result)        if result.verdict != "SUPPORTED":            return {"verdict": "UNVERIFIED", "failed_at": step["name"], "trace": trace}        context[step["binds"]] = step["extract"](result)    return {"verdict": "SUPPORTED", "trace": trace}

The failed_at field is the part worth keeping. "This claim failed at hop 2 because we could not establish who acquired the supplier" is actionable; "unverified" is not.

What this means when you run grounded generation in production

Instrument the retriever separately from the verifier, and put both on the same chart. Recall@k on a frozen labelled evidence set is a leading indicator: when it drops after a corpus re-index or an embedding-model upgrade, your faithfulness numbers will drop a week later and look like a model regression. Teams without retriever metrics reliably spend that week debugging the generator.

Return evidence identifiers with every verdict and store them. A faithfulness rate you cannot drill into is a number you cannot act on; a stored evidence trace lets you answer "why did we mark this unsupported" months later, when someone challenges a flagged response. It also gives you the labelled data for your next threshold sweep, essentially for free.

Finally, decide what the system does when a claim fails, before you build the checker. The options are: block the response, strip the unsupported sentence, append a hedge, or log and pass through. Each is right somewhere, and each implies a different threshold. A checker built without that decision made tends to end up as a dashboard nobody acts on — measuring, precisely and continuously, a problem that nothing in the system is empowered to fix.