RAG Systems

Course Content

RAG Systems

12 sections · 66 lessons

What are common failure modes in RAG systems?


What you need to know

A useful reference is the 2024 paper Seven Failure Points When Engineering a Retrieval Augmented Generation System (Barnett and others). Its seven points map neatly onto pipeline stages:

Failure pointWhat happensStage
Missing contentThe answer is not in the corpus, and the system answers anywayIngestion / generation
Missed top-ranked documentsThe answer exists but is not ranked high enoughRetrieval / ranking
Not in contextRetrieved, but dropped when building the promptPrompt building
Not extractedIn the prompt, but the model fails to use itGeneration
Wrong formatAsked for a table or list, got something elseGeneration
Incorrect specificityToo general or too detailed for the questionGeneration
IncompletePart of the answer missing, though availableGeneration

Stage by stage, with the fix

  • Ingestion. Parser drops tables, OCR fails on scans, a connector skips attachments, jobs re-run with random ids. Fix: parse checks, count reconciliation, deterministic ids.
  • Chunking. A limit value is split from the row saying which grade it applies to. Fix: structure-aware splitting, keep tables whole, add section headings or context to chunks.
  • Retrieval. Dense embeddings blur rare tokens such as order ids, error codes and SKUs. Follow-up questions like "what about the blue one?" have no meaning alone. Fix: hybrid search with BM25, query rewriting, checking filters.
  • Ranking. Right chunk at rank 18 when you keep 5. Fix: reranker, better first-stage recall.
  • Generation. Ignoring context, answering from training data, refusing answerable questions. Fix: grounding prompt, citation and faithfulness checks, tuning the relevance floor.
  • Operations. Index stale after a source update, embedding model upgraded for queries but not for the index, tenant filter missing in one code path. Fix: freshness monitoring, versioned indexes, access-control tests.
  • Security. Instructions hidden in a retrieved document change the model's behaviour (indirect prompt injection). Fix: covered in the security section.

A real-life example

An e-commerce product Q&A bot is reviewed after a sale weekend. The team samples 100 answers with negative feedback and labels the first failing stage from each trace:

  • 34: the answer used last season's price or stock text (stale index; the price feed updated but the Q&A index re-ingested weekly).
  • 21: the product code in the question was not matched (dense-only search).
  • 17: the right chunk was found but ranked below the top 4.
  • 14: the model added claims from reviews as facts.
  • 9: the question had no answer in the catalogue, and the bot answered anyway.
  • 5: other.

The biggest category is operational, not the model. The team moves price and stock out of RAG entirely: they are fetched live from the product API as a tool call, while descriptions stay in the index. Then they add BM25 and a reranker for the next two groups.

Follow-up questions to expect

  • "How do you tell a retrieval failure from a generation failure?" — Look at the retrieved chunks. If the answer is not there, it is retrieval (or missing content). If it is there, it is generation.
  • "How do you handle questions with no answer in the corpus?" — A relevance floor that returns a clear "not found", plus test cases with no answer in the golden set.
  • "Which failure is most common?" — In most production reviews, data problems: stale, duplicated, missing or badly parsed documents.