Course Content
Scenario-Based AI Engineering Questions
26 sections · 146 lessons
Scenario – 9: Hallucination Control
What you need to know
The scenario: a research crew's report contains a confident market-size figure that no source supports.
Why crews make it worse
Each agent trusts the agent before it. An invented figure in the first task becomes an "input fact" for the second, and a "finding" in the final report. The more agents repeat it, the more established it looks, which is why this is called laundering.
Controls that stop propagation
- Tools for facts — any agent asked for facts gets a search or retrieval tool. A "researcher" without one is a fiction generator with a job title.
- Sources required — each claim is an object with
textandsource_id; a guardrail rejects claims without one. - Verify before assembly — a fact-check task checks each claim against its cited source and returns a verdict per claim.
- Copy, don't retype — downstream tasks receive numbers as typed fields; the writer formats them and never re-derives them.
- Allow "unknown" —
Noneplus a reason is better than a confident guess; use temperature 0 for factual tasks.
1class Claim(BaseModel):2 text: str3 value: float | None = None4 source_id: str | None5 quote: str | None # the exact supporting sentence from the source67class Research(BaseModel):8 claims: list[Claim]910def sourced(result: TaskOutput):11 bad = [c.text for c in result.pydantic.claims if not (c.source_id and c.quote)]12 if bad:13 return (False, f"These claims need a source and quote, or remove them: {bad[:3]}")14 return (True, result)Asking for the exact supporting quote makes checking easy: the verifier confirms the quote exists in the source and entails the claim.
Metrics per task
| Metric | Target |
|---|---|
| Unsourced-claim rate | Zero, by construction |
| Citation precision on a sampled audit | High; the cited source really supports the claim |
| End-to-end factual accuracy on a labelled set | Tracked per task, because losses concentrate in one agent |
A real-life example
Scenario, numbers made up. A consulting crew's report on India's EV two-wheeler market states a market size that no source supports; the analyst built a five-year forecast on it and the writer made it the headline. A partner catches it before the client does.
The team gives the researcher web and database search, requires source_id and quote for every claim, adds a verification task with an NLI check plus an LLM judge for flagged claims, and passes numbers to the writer as fields. On a 30-report audit, unsupported claims fall from 12% to under 1%, and reports now say "no reliable figure found for 2023 sales in tier-3 cities" instead of inventing one.
Follow-up questions to expect
- "Isn't a fact-check agent just another LLM that can be wrong?" — Yes, so it checks quotes against sources mechanically first and uses a judge only for the harder cases, with human audits on a sample.
- "What about the model's own knowledge?" — For client-facing facts, if it cannot be sourced, it is labelled as unverified or removed.
- "Where do you check first?" — The earliest task that produces facts; stopping an error there prevents every later agent building on it.