Course Content
RAG Systems
12 sections · 66 lessons
What are common failure modes in RAG systems?
What you need to know
A useful reference is the 2024 paper Seven Failure Points When Engineering a Retrieval Augmented Generation System (Barnett and others). Its seven points map neatly onto pipeline stages:
| Failure point | What happens | Stage |
|---|---|---|
| Missing content | The answer is not in the corpus, and the system answers anyway | Ingestion / generation |
| Missed top-ranked documents | The answer exists but is not ranked high enough | Retrieval / ranking |
| Not in context | Retrieved, but dropped when building the prompt | Prompt building |
| Not extracted | In the prompt, but the model fails to use it | Generation |
| Wrong format | Asked for a table or list, got something else | Generation |
| Incorrect specificity | Too general or too detailed for the question | Generation |
| Incomplete | Part of the answer missing, though available | Generation |
Stage by stage, with the fix
- Ingestion. Parser drops tables, OCR fails on scans, a connector skips attachments, jobs re-run with random ids. Fix: parse checks, count reconciliation, deterministic ids.
- Chunking. A limit value is split from the row saying which grade it applies to. Fix: structure-aware splitting, keep tables whole, add section headings or context to chunks.
- Retrieval. Dense embeddings blur rare tokens such as order ids, error codes and SKUs. Follow-up questions like "what about the blue one?" have no meaning alone. Fix: hybrid search with BM25, query rewriting, checking filters.
- Ranking. Right chunk at rank 18 when you keep 5. Fix: reranker, better first-stage recall.
- Generation. Ignoring context, answering from training data, refusing answerable questions. Fix: grounding prompt, citation and faithfulness checks, tuning the relevance floor.
- Operations. Index stale after a source update, embedding model upgraded for queries but not for the index, tenant filter missing in one code path. Fix: freshness monitoring, versioned indexes, access-control tests.
- Security. Instructions hidden in a retrieved document change the model's behaviour (indirect prompt injection). Fix: covered in the security section.
A real-life example
An e-commerce product Q&A bot is reviewed after a sale weekend. The team samples 100 answers with negative feedback and labels the first failing stage from each trace:
- 34: the answer used last season's price or stock text (stale index; the price feed updated but the Q&A index re-ingested weekly).
- 21: the product code in the question was not matched (dense-only search).
- 17: the right chunk was found but ranked below the top 4.
- 14: the model added claims from reviews as facts.
- 9: the question had no answer in the catalogue, and the bot answered anyway.
- 5: other.
The biggest category is operational, not the model. The team moves price and stock out of RAG entirely: they are fetched live from the product API as a tool call, while descriptions stay in the index. Then they add BM25 and a reranker for the next two groups.
Follow-up questions to expect
- "How do you tell a retrieval failure from a generation failure?" — Look at the retrieved chunks. If the answer is not there, it is retrieval (or missing content). If it is there, it is generation.
- "How do you handle questions with no answer in the corpus?" — A relevance floor that returns a clear "not found", plus test cases with no answer in the golden set.
- "Which failure is most common?" — In most production reviews, data problems: stale, duplicated, missing or badly parsed documents.