LLM Evaluation

Course Content

LLM Evaluation

6 sections · 50 lessons

How does search-based verification improve factual evaluation?


What you need to know

The pipeline

Python
def verify(answer: str) -> list[dict]:    results = []    for claim in decompose(answer):              # LLM splits into atomic facts        query = make_query(claim)                 # rewrite claim as a search query        evidence = search(query, k=5)             # web, or your own knowledge base        verdict = judge_entailment(evidence, claim)  # supported / refuted / no_evidence        results.append({"claim": claim, "verdict": verdict, "evidence": evidence})    return results

decompose, search and judge_entailment are your own functions — an LLM call, a search API and an NLI model or judge. Each claim ends with a verdict and the evidence behind it.

The metric

FActScore (2023) defines factual precision as the share of atomic facts supported by a knowledge source (Wikipedia in the paper). SAFE (2024) replaces a fixed source with multi-step Google Search queries by an LLM agent, and reports that it agreed with crowd-sourced human raters most of the time at much lower cost.

Text
factual precision = supported claims / (supported + refuted claims)report separately:  no-evidence claims

Keep "no evidence" separate. Mixing it into "false" punishes the system for search failures.

Why it beats self-consistency

It can catch errors the model makes every time, because the evidence comes from outside the model. And it leaves an audit trail: each verdict has a source link a human can open.

Trade-offs

  • Cost and latency — a search plus a judge call per claim; an answer with 12 claims means 24 or more calls.
  • Query formulation — a bad query returns nothing relevant, and the claim looks unverifiable.
  • Source quality — the web has outdated and conflicting pages. For company facts, search your own documents, not the web.
  • Private or subjective claims — "our refund SLA is 5 days" cannot be checked on the web; "this phone is stylish" cannot be checked at all.

A real-life example

The e-commerce team verifies product descriptions for its electronics category against a trusted source: the manufacturer spec pages they have licensed and indexed. For each description, a model extracts claims (battery, weight, water rating, warranty), searches the index, and runs an entailment check.

On 1,000 descriptions with about 9,000 claims: 93% supported, 2% refuted, 5% no evidence. The refuted claims are almost all wrong numbers — "IP68" where the spec says IP54 — the kind of error SelfCheckGPT missed because the model repeated it every time. The no-evidence group is mostly marketing phrases ("perfect for gaming"), which the team decides to allow but not to write as facts. Refuted claims block publication; each has a link to the spec page, so sellers accept the correction.

Follow-up questions to expect

  • "How do you evaluate the verifier itself?" — Label a few hundred claims with verdicts by hand and measure the pipeline's precision and recall for "refuted", and how often "no evidence" was a search failure.
  • "When would you search the web versus an internal index?" — Web for public facts, internal index for company-specific facts; never the open web for private policies.
  • "Can this run on every request?" — Usually not; run it offline, on high-risk outputs, or on a sample. For real-time, check only claim types that matter, like numbers.