RAG Systems

Course Content

RAG Systems

12 sections · 66 lessons

What are the limitations of a pure LLM system without RAG?


What you need to know

Think of a pure model as a very well-read person locked in a room since a certain date, with no phone and no files. Their reasoning is good; their facts are fixed.

The six limitations

LimitationWhat it meansWhat it looks like in production
Knowledge cutoffTraining stopped at a dateQuotes last year's interest rate
No private dataYour wiki, contracts and tickets were never in trainingInvents a leave policy that sounds typical
HallucinationIt fills gaps with likely-sounding textConfident answer, wrong fact, same tone as a right one
No provenanceNo source to point toCompliance team cannot audit the answer
Costly updatesFacts live in weightsFixing one wrong fact needs a training run
No access controlWeights cannot be filtered per userA contractor could get an answer drawn from board papers

The dangerous part is the third row. A wrong answer and a right answer look identical to the user. Nothing in the text says "I am guessing".

Why "just paste everything in the prompt" is not a full fix

Long context windows help for small document sets, but they have costs:

  • Money. Every token you send is billed, on every call. 200,000 tokens of policy documents sent 10,000 times a day is two billion input tokens a day. Prompt caching lowers this but does not remove it.
  • Latency. The model must read the whole prompt before it writes the first word.
  • Accuracy. Studies such as "Lost in the Middle" (Liu et al., 2023) showed models use facts in the middle of long prompts less reliably than facts near the start or end. Newer models are better, but quality still tends to drop as prompts grow.
  • Security. One shared prompt cannot give different users different documents.

A real-life example

An Indian private bank launches a product-FAQ chatbot on a plain model. A customer asks: "What is the interest rate on a 1-year fixed deposit for senior citizens?"

The model answers "7.25%". That number appeared on many web pages during training. The bank changed its rate to 7.10% last Tuesday. The customer opens a deposit expecting 7.25%, complains, and the complaint reaches the compliance team, who ask "where did the bot get this number?". There is no answer, because the number came from the weights.

The fix has two parts. Rates come from the bank's rate API at query time, because they are live data. Product rules (premature withdrawal penalty, minimum amount, tax deduction rules) come from retrieved, versioned product documents, with a citation in every reply. Now each answer can be traced to a source the bank controls.

Follow-up questions to expect

  • "Won't a bigger or newer model fix this?" — It moves the cutoff and reduces some errors, but it still has never seen your private data and still cannot cite or be filtered per user.
  • "Can the model tell when it does not know?" — Only partly. Reasoning models refuse more often, but you cannot rely on it. Giving the model evidence and permission to say "not found" is the dependable fix.
  • "Is web search a form of RAG?" — Yes. A search call that returns pages which the model then reads is retrieval-augmented generation with the web as the index.