Course Content
RAG Systems
12 sections · 66 lessons
What is Retrieval Augmented Generation (RAG), and why is it critical for LLMs?
What you need to know
A language model stores knowledge in two places. Parametric memory is what it learned during training, spread across its weights. You cannot edit it, cite it or check it. Non-parametric memory is text you hand the model in the prompt. You can edit it, delete it, filter it per user and point to its source. RAG moves the facts from the first kind to the second.
The two halves
- Index (offline) — load documents, split them into chunks of a few hundred tokens, turn each chunk into an embedding vector, and store vectors, text and metadata in a search index.
- Retrieve (per request) — turn the user's question into a query, search the index, and keep the top few chunks.
- Augment — place those chunks in the prompt, each with a source label, plus an instruction to answer only from them.
- Generate — the model writes the answer and cites the chunk labels it used.
Retrieval does not have to be a vector database. It can be keyword search (BM25), a SQL query, a web search or a mix. "Vector DB plus LLM" is one common way to build RAG, not the definition of it.
Why it still matters when context windows are huge
In 2026 several models accept around a million tokens of input. That changes when you need RAG, not whether:
- Small, stable corpus. A 60-page product manual is about 40,000 tokens. You can put all of it in every prompt and use prompt caching to cut the repeat cost. RAG may be unnecessary here.
- Large or changing corpus. A bank's 30,000 documents are tens of millions of tokens. No window holds that, and re-sending even a tenth of it on every question is slow and expensive.
- Access control. An employee should only see documents they are allowed to see. Retrieval with filters enforces that; one giant prompt cannot.
- Quality. Models use information in very long prompts less reliably, especially material far from the start and end. Sending the right 3,000 tokens often beats sending 300,000.
A real-life example
A company with 5,000 employees builds an HR policy assistant. An employee asks: "I joined two years ago. How many days of paternity leave can I take?"
Without RAG, the model gives a plausible general answer, such as "most companies offer 5 to 15 days". It sounds right and is useless, because this company's policy is its own.
With RAG, the retriever finds a 350-token chunk from Leave Policy v4.2, section 3.4: "Employees with more than one year of service are eligible for 10 working days of paternity leave, to be taken within 6 months of the child's birth." The assistant answers "10 working days, within 6 months of the birth [Leave Policy v4.2, §3.4]".
When HR changes the policy to 15 days next quarter, the team re-indexes one document. That takes seconds. Nobody retrains anything.
Follow-up questions to expect
- "Does RAG remove hallucinations?" — No, it reduces them. The model can still misread a chunk or ignore it, and a retrieval miss can still lead to a guess. You measure faithfulness (is every claim supported by a retrieved passage) and allow the model to say "not in the documents".
- "Why not just use a million-token context?" — For a small, static corpus I would. For a large, changing or permissioned corpus, RAG is cheaper, faster and can enforce who sees what.
- "Is RAG only for question answering?" — No. The same pattern grounds summarisation, drafting emails from a CRM, code assistants reading a repository, and agents that search before acting.