Course Content
AutoGen Essentials
7 sections · 28 lessons
How do you detect and reduce hallucinations in agent-produced decisions and tool arguments?
What you need to know
Two kinds, two sets of fixes
| Where | Example | Harm | Main fixes |
|---|---|---|---|
| Tool arguments | cancel_booking(pnr="4521987654") for a PNR the user never gave | A real wrong side effect | Schemas, existence checks, grounded ids |
| Answers and decisions | "Your flight is refundable" when the fare rules say not | Wrong advice, broken trust | Citations, "insufficient information" path, verifier |
Fixing tool arguments
- Types and enums. AutoGen validates arguments against your type hints. Use
Literal["economy", "business"]or Pydantic models so wrong values fail before the side effect. - Existence checks. "PNR 4521987654 not found for this customer" is far better than an empty result the model will explain away.
- Grounded ids. Only accept ids that appeared in the user's messages or an earlier tool result:
1import re23def grounded(value: str, messages) -> bool:4 seen = " ".join(m.to_text() for m in messages)5 return re.search(rf"\b{re.escape(value)}\b", seen) is not None67async def cancel_booking(pnr: str) -> dict:8 """Cancel a booking by its 10-digit PNR from the user or a lookup."""9 if not grounded(pnr, session.messages):10 return {"ok": False, "error": "ungrounded_pnr",11 "hint": "Ask the user for the PNR or look it up first."}12 ...- Identity server-side. Never let the model pass
user_idortenant_id; inject them from the session.
Fixing answers and decisions
- Citations that are checked. Each claim points to a retrieved chunk or tool result, and code or a verifier checks the chunk supports it.
- An honest exit. "If the evidence does not answer the question, say so and say what is missing." Without this path, models fill the gap.
- A separate verifier. An agent that sees the evidence and the answer, but not the writer's reasoning, catches more unsupported claims than self-review.
- Lower the pressure. Shorter context, fewer tools per agent and a clear task all reduce the rate.
Measure it
Sample runs each week; label each claim supported, unsupported or contradicted; track the unsupported rate per release like an error rate.
A real-life example
An airline's travel-planning assistant handled changes and cancellations. In a month of logs, 1.8% of cancel_booking calls used a PNR that did not appear anywhere in the conversation; the model had "completed" a partial number. Two actually matched other passengers' bookings, which the backend luckily rejected because the name did not match.
The team added the grounding check, a strict 10-digit pattern type, and an existence-plus-ownership check in the tool. Ungrounded calls now return an error, and the agent asks the user. Separately, a verifier agent now checks any "refundable" or "free change" statement against the fare-rule text returned by get_fare_rules. Unsupported claims about refunds fell from 6% of sampled answers to under 1%.
Follow-up questions to expect
- "Can't you just tell the model not to hallucinate?" — A prompt helps a little; structure helps much more: validation, grounding checks and verifiers that do not depend on the model obeying.
- "How does a verifier differ from a critic?" — A verifier checks facts against evidence with a narrow yes/no job; a critic judges overall quality. Verifiers are easier to make reliable.
- "Does lower temperature fix it?" — It reduces randomness, not wrong knowledge. A model can be confidently and consistently wrong at temperature 0.