Course Content
LLM Evaluation
6 sections · 50 lessons
How does search-based verification improve factual evaluation?
What you need to know
The pipeline
1def verify(answer: str) -> list[dict]:2 results = []3 for claim in decompose(answer): # LLM splits into atomic facts4 query = make_query(claim) # rewrite claim as a search query5 evidence = search(query, k=5) # web, or your own knowledge base6 verdict = judge_entailment(evidence, claim) # supported / refuted / no_evidence7 results.append({"claim": claim, "verdict": verdict, "evidence": evidence})8 return resultsdecompose, search and judge_entailment are your own functions — an LLM call, a search API and an NLI model or judge. Each claim ends with a verdict and the evidence behind it.
The metric
FActScore (2023) defines factual precision as the share of atomic facts supported by a knowledge source (Wikipedia in the paper). SAFE (2024) replaces a fixed source with multi-step Google Search queries by an LLM agent, and reports that it agreed with crowd-sourced human raters most of the time at much lower cost.
factual precision = supported claims / (supported + refuted claims)report separately: no-evidence claimsKeep "no evidence" separate. Mixing it into "false" punishes the system for search failures.
Why it beats self-consistency
It can catch errors the model makes every time, because the evidence comes from outside the model. And it leaves an audit trail: each verdict has a source link a human can open.
Trade-offs
- Cost and latency — a search plus a judge call per claim; an answer with 12 claims means 24 or more calls.
- Query formulation — a bad query returns nothing relevant, and the claim looks unverifiable.
- Source quality — the web has outdated and conflicting pages. For company facts, search your own documents, not the web.
- Private or subjective claims — "our refund SLA is 5 days" cannot be checked on the web; "this phone is stylish" cannot be checked at all.
A real-life example
The e-commerce team verifies product descriptions for its electronics category against a trusted source: the manufacturer spec pages they have licensed and indexed. For each description, a model extracts claims (battery, weight, water rating, warranty), searches the index, and runs an entailment check.
On 1,000 descriptions with about 9,000 claims: 93% supported, 2% refuted, 5% no evidence. The refuted claims are almost all wrong numbers — "IP68" where the spec says IP54 — the kind of error SelfCheckGPT missed because the model repeated it every time. The no-evidence group is mostly marketing phrases ("perfect for gaming"), which the team decides to allow but not to write as facts. Refuted claims block publication; each has a link to the spec page, so sellers accept the correction.
Follow-up questions to expect
- "How do you evaluate the verifier itself?" — Label a few hundred claims with verdicts by hand and measure the pipeline's precision and recall for "refuted", and how often "no evidence" was a search failure.
- "When would you search the web versus an internal index?" — Web for public facts, internal index for company-specific facts; never the open web for private policies.
- "Can this run on every request?" — Usually not; run it offline, on high-risk outputs, or on a sample. For real-time, check only claim types that matter, like numbers.