Course Content
LLMs Deep Dive
10 sections · 40 lessons
An LLM generates offensive or incorrect outputs. How do you handle it?
What you need to know
Why no single fix is enough
An LLM's output depends on the model, the prompt, the retrieved context, the user's input and the sampling settings. A bad output can start in any of them, and the model itself has no reliable truth check. So defences go around the model, not only inside the prompt.
The response procedure
- Contain — if users are being harmed, disable the feature, switch to a fallback message, or roll back the prompt or model version.
- Capture — log the full request: user input, system prompt version, retrieved documents, model name and version, settings, and the output.
- Reproduce — replay the request; run it several times, since sampling varies.
- Diagnose — decide which layer failed: prompt, retrieval, injection, model change or out-of-scope use.
- Fix in layers — see below.
- Prevent recurrence — add the case to the regression suite, monitor, and write a short incident note.
The layers of defence
- Input — classify and block disallowed requests; mark retrieved and user-pasted text as untrusted data, not instructions.
- Prompt — clear scope, refusal rules and "answer only from the provided context; otherwise say you don't know".
- Retrieval — correct, current sources; require citations so claims can be checked.
- Output — a moderation classifier; schema validation; code checks on facts that matter (amounts, dates, product specs) against the source system.
- People — human approval for high-impact actions such as refunds or legal advice; an easy "report this answer" button.
- Monitoring — track flag rates, user reports and regression-test scores per release.
A real-life example
A bank's support bot tells a customer in Hinglish, "Aapke home loan ka pre-closure bilkul free hai" ("pre-closure of your home loan is completely free"). The customer's loan actually carries a 2% pre-closure charge. Separately, a user pastes a "complaint letter" containing hidden instructions, and the bot replies with an insulting line.
The team disables pre-closure answers and shows "please speak to an agent" within the hour. Logs show two different causes. For the first, retrieval returned the policy for floating-rate loans (no charge) instead of fixed-rate loans (2% charge); the fix adds loan type as a retrieval filter and a check that any charge the bot quotes matches the loan record. For the second, the pasted letter was treated as instructions; the fix wraps user-pasted text as data, adds an output moderation check, and adds 30 injection examples to the regression suite. Both cases join a 500-case test set that every prompt or model change must pass.
Follow-up questions to expect
- "How do you find these failures before users do?" — A regression suite including adversarial cases, red-teaming before launch, and sampling live conversations for review.
- "Would a stronger model solve it?" — It may reduce the rate, but not to zero; wrong retrieval and prompt injection affect every model, so the layers are still needed.
- "Who owns the incident?" — The product team, with safety, legal and support involved for serious cases; the key is a clear rollback path decided in advance.