Course Content
Live Coding Interview Prep
7 sections · 50 lessons
Implement hallucination detection using multiple responses.
What you need to know
A hallucination is a fluent statement that is not true or not supported by the given sources. With retrieved context, you check the answer against the context. Without a source, you need another signal — and one of the best is consistency.
The intuition: ask a friend the capital of Australia five times and you get "Canberra" five times. Ask them the name of your neighbour's cousin and, if they are bluffing, you get five different names. Sampling a model at temperature above 0 is asking it several times.
How to decide whether a sample "supports" a sentence:
| method | how | quality |
|---|---|---|
| Embedding similarity | cosine between the sentence and the closest sentence in the sample | cheap; weak on numbers and negation |
| NLI model | "does the sample entail this sentence?" | much better; one small model call per pair |
| LLM question | "is this sentence supported by this passage? yes/no" | best; most expensive |
The limitation to state: self-consistency cannot catch a hallucination the model makes consistently. If it always says the wrong date, all samples agree. That is why checking against a source is better whenever you have one.
1import re2from collections.abc import Callable3import numpy as np45def split_sentences(text: str) -> list[str]:6 return [s.strip() for s in re.split(r"(?<=[.!?])\s+", text) if s.strip()]78def _unit(embed_fn: Callable, texts: list[str]) -> np.ndarray:9 v = np.asarray(embed_fn(texts), dtype=np.float32)10 return v / (np.linalg.norm(v, axis=1, keepdims=True) + 1e-10)1112def self_check(primary: str, samples: list[str], embed_fn: Callable,13 sim_threshold: float = 0.8, support_threshold: float = 0.5) -> dict:14 """Flag sentences of `primary` that fewer than half the samples support."""15 sentences = split_sentences(primary)16 per_sample = [split_sentences(s) for s in samples]17 per_sample = [ss for ss in per_sample if ss]18 if not sentences or not per_sample:19 return {"support": [], "flagged": sentences}20 a = _unit(embed_fn, sentences)21 support = np.zeros(len(sentences))22 for ss in per_sample: # does this sample agree?23 best = (a @ _unit(embed_fn, ss).T).max(axis=1) # closest sentence in the sample24 support += best >= sim_threshold25 support /= len(per_sample)26 flagged = [s for s, sup in zip(sentences, support) if sup < support_threshold]27 return {"support": [round(float(x), 2) for x in support], "flagged": flagged}The function takes the samples as input, so generation stays outside it: primary = generate(prompt, temperature=0) and samples = [generate(prompt, temperature=1) for _ in range(k)].
The tricky parts:
(a @ B.T).max(axis=1)— a matrix of similarities between every answer sentence and every sentence of one sample, then the best match per answer sentence. A sentence is supported by a sample if any of its sentences says the same thing.support += best >= sim_thresholdadds 1 for each sample that supports each sentence (booleans add as 0 or 1). Dividing by the number of samples gives a fraction.- Empty samples are skipped; with no usable samples, everything is flagged rather than silently passed.
Complexity: k + 1 generations dominate. Embedding is one call for the answer plus one per sample. The similarity work is O(n·m·d) for n answer sentences and m sample sentences in total. Space O(n·m) for one sample's matrix at a time.
A real-life example
The answer has one true and one false sentence (Tagore was born in Kolkata, not Mumbai). Hand-made 3-dimensional embeddings: the first number means "Nobel in 1913", the second "born in Mumbai", the third "born in Kolkata":
1VECS = {"Tagore won the Nobel Prize in 1913.": [1, 0, 0],2 "He was born in Mumbai.": [0, 1, 0],3 "In 1913 Tagore received the Nobel Prize.": [0.95, 0, 0.1],4 "He was born in Kolkata.": [0, 0.2, 1],5 "Tagore got the Nobel in 1913.": [0.9, 0.1, 0],6 "His birthplace was Calcutta.": [0, 0.1, 1],7 "The Nobel came to Tagore in 1913.": [1, 0, 0.05],8 "Tagore was born in Mumbai.": [0, 1, 0.1]}9embed = lambda ts: [VECS[t] for t in ts]10primary = "Tagore won the Nobel Prize in 1913. He was born in Mumbai."11samples = ["In 1913 Tagore received the Nobel Prize. He was born in Kolkata.",12 "Tagore got the Nobel in 1913. His birthplace was Calcutta.",13 "The Nobel came to Tagore in 1913. Tagore was born in Mumbai."]14print(self_check(primary, samples, embed))15# {'support': [1.0, 0.33], 'flagged': ['He was born in Mumbai.']}| answer sentence | sample 1 | sample 2 | sample 3 | support |
|---|---|---|---|---|
| Nobel Prize in 1913 | 0.99 ✓ | 0.99 ✓ | 1.00 ✓ | 3/3 = 1.0 |
| born in Mumbai | 0.20 ✗ | 0.10 ✗ | 0.99 ✓ | 1/3 = 0.33 → flagged |
The Nobel sentence is stable across samples; the birthplace flips, so it is flagged. Notice sample 3 repeated the error — if all three had, the check would have passed it.
An encyclopaedia-style Q&A feature, or a tool that drafts biographies for a news site, can run this offline on a sample of answers to measure how often the model makes things up.
Follow-up questions to expect
- "Why not just ask the model if it is sure?" — Verbal confidence is poorly calibrated; the model sounds equally sure when right and wrong. Consistency across samples is a behavioural signal, not a self-report.
- "What if you have retrieved documents?" — Check each sentence against the documents with an NLI model or a small LLM instead. It is cheaper (no extra generations) and catches consistent errors.
- "How do you evaluate the detector?" — Build a set with known hallucinations, for example questions about invented people, and report precision and recall of flagging at several thresholds.