RAG Systems

Course Content

RAG Systems

12 sections · 66 lessons

What types of problems are best solved using RAG?


Which mechanism answers which questionQuestion routerAnswer in a fewpassages: RAGCount or filtercontracts: SQLOrder status or balance: APIArithmetic: calculator toolThemes acrosseverything: map-reduce
Top-k retrieval returns a handful of passages, so a question needing every document gets a confident partial answer.

What you need to know

Top-k retrieval returns a handful of passages, typically 3 to 10. So the honest test is: can the answer be written from 3 to 10 passages? If yes, RAG fits. If the answer needs every document, or a calculation, or live data, it does not.

Good fits and poor fits

Good fit for RAG

  • "What is the notice period in my contract?"
  • "How do I reset the router to factory settings?"
  • "What did we decide in the March incident review?"
  • "Which API endpoint returns the refund status?"

Poor fit for plain RAG

  • "How many contracts expire in Q3?" (aggregation)
  • "What is my account balance?" (live state)
  • "What is the average claim value by city?" (computation)
  • "Summarise every complaint from last year" (whole corpus)

What to use instead for the poor fits

  • Aggregation and filters — extract structured fields at ingest time (expiry date, party, value) into a table, then answer with SQL. An LLM can write the SQL (text-to-SQL).
  • Live state — call the system of record through an API or a tool.
  • Computation — let the model call a calculator or run code.
  • Whole-corpus questions — map-reduce summarisation (summarise each part, then combine), or a knowledge-graph approach such as GraphRAG that builds community summaries in advance.

Many real assistants route between these. A router, often an LLM call with tool definitions, decides whether a question needs retrieval, SQL, an API or nothing at all.

A real-life example

A legal-contract search tool at a mid-size company holds 8,000 vendor contracts. Two questions arrive on the same day.

  1. "What does our contract with a logistics vendor say about liability for damaged goods?" The answer is in one or two clauses. Hybrid search finds the indemnity and limitation-of-liability sections, and the model quotes them with page numbers. This is classic RAG.
  2. "Which contracts auto-renew with a notice period under 60 days?" Top-k retrieval returns 10 chunks, so the answer would cover 10 contracts, not all of them. The model would report a confident but partial list. The team adds an extraction job: at ingest, an LLM reads each contract once and fills a table with auto_renews, notice_days and renewal_date. The question becomes SELECT ... WHERE auto_renews AND notice_days < 60, and the answer is complete.

The same product serves both, but with two different mechanisms behind one chat box.

Follow-up questions to expect

  • "How would you route between RAG and SQL?" — Give the model two tools, search_documents and query_contracts_table, with clear descriptions, and let it choose. Log the choices and evaluate routing accuracy on labelled questions.
  • "What is GraphRAG?" — An approach that extracts entities and relationships into a graph and pre-computes summaries of clusters, so it can answer broad "what are the main themes" questions that top-k search cannot.
  • "Can RAG handle multi-hop questions?" — Only with several retrieval steps: find the first fact, use it to form the next query. That is covered under agentic RAG.