Course Content
RAG Systems
12 sections · 66 lessons
What types of problems are best solved using RAG?
What you need to know
Top-k retrieval returns a handful of passages, typically 3 to 10. So the honest test is: can the answer be written from 3 to 10 passages? If yes, RAG fits. If the answer needs every document, or a calculation, or live data, it does not.
Good fits and poor fits
Good fit for RAG
- "What is the notice period in my contract?"
- "How do I reset the router to factory settings?"
- "What did we decide in the March incident review?"
- "Which API endpoint returns the refund status?"
Poor fit for plain RAG
- "How many contracts expire in Q3?" (aggregation)
- "What is my account balance?" (live state)
- "What is the average claim value by city?" (computation)
- "Summarise every complaint from last year" (whole corpus)
What to use instead for the poor fits
- Aggregation and filters — extract structured fields at ingest time (expiry date, party, value) into a table, then answer with SQL. An LLM can write the SQL (text-to-SQL).
- Live state — call the system of record through an API or a tool.
- Computation — let the model call a calculator or run code.
- Whole-corpus questions — map-reduce summarisation (summarise each part, then combine), or a knowledge-graph approach such as GraphRAG that builds community summaries in advance.
Many real assistants route between these. A router, often an LLM call with tool definitions, decides whether a question needs retrieval, SQL, an API or nothing at all.
A real-life example
A legal-contract search tool at a mid-size company holds 8,000 vendor contracts. Two questions arrive on the same day.
- "What does our contract with a logistics vendor say about liability for damaged goods?" The answer is in one or two clauses. Hybrid search finds the indemnity and limitation-of-liability sections, and the model quotes them with page numbers. This is classic RAG.
- "Which contracts auto-renew with a notice period under 60 days?" Top-k retrieval returns 10 chunks, so the answer would cover 10 contracts, not all of them. The model would report a confident but partial list. The team adds an extraction job: at ingest, an LLM reads each contract once and fills a table with
auto_renews,notice_daysandrenewal_date. The question becomesSELECT ... WHERE auto_renews AND notice_days < 60, and the answer is complete.
The same product serves both, but with two different mechanisms behind one chat box.
Follow-up questions to expect
- "How would you route between RAG and SQL?" — Give the model two tools,
search_documentsandquery_contracts_table, with clear descriptions, and let it choose. Log the choices and evaluate routing accuracy on labelled questions. - "What is GraphRAG?" — An approach that extracts entities and relationships into a graph and pre-computes summaries of clusters, so it can answer broad "what are the main themes" questions that top-k search cannot.
- "Can RAG handle multi-hop questions?" — Only with several retrieval steps: find the first fact, use it to form the next query. That is covered under agentic RAG.