Course Content
RAG Systems
12 sections · 66 lessons
What are the two main components of a RAG application architecture?
What you need to know
The split exists because the two halves have completely different needs.
Indexing (offline)
- Runs on a schedule or on document changes
- Throughput matters, latency does not
- Heavy work: parsing, OCR, embedding millions of chunks
- Output: a versioned search index
Retrieval and generation (online)
- Runs once per user question
- Latency matters: users wait for it
- Light work per call: one query embedding, one search, one LLM call
- Output: an answer with citations
Why this matters when you debug
A wrong answer appears at request time, but its cause is often upstream:
- A PDF table was flattened into loose numbers during parsing.
- A heading was split away from its paragraph during chunking.
- The index was built with a different embedding model from the one used for queries.
- A document was updated, but the old chunks were never deleted.
None of these can be fixed with a better prompt. That is why experienced engineers look at the retrieved chunks, and then at the index, before touching the prompt or the model.
Where the pieces run
| Component | Typical home |
|---|---|
| Ingestion workers | A queue plus batch jobs (Airflow, a cron job, a serverless function) |
| Embedding model | A hosted API or a GPU service you run |
| Index | A vector database (Qdrant, Weaviate, Milvus, Pinecone) or Postgres with pgvector, often alongside a keyword index |
| Request path | Your API service: retriever, reranker, prompt builder, LLM client |
| Logs and traces | A tracing tool, so every answer can be replayed with its retrieved chunks |
You do not always need a dedicated vector database. Up to a few million chunks, pgvector inside a Postgres you already run is often the simplest choice.
A real-life example
An HR policy assistant goes live and, within a week, employees report that it says the notice period for band 4 is "60 days". The policy says 90.
The team first looks at the prompt, which seems fine. Then they print the retrieved chunks for that question. The top chunk ends with "Employees in bands 1 to 3 must serve a notice period of 60 days." The next sentence, the one about band 4, is in a different chunk that ranked sixth, outside the top 4.
The bug was created at indexing time: a 100-token chunk size split one rule into two chunks. They re-index with structure-aware chunking so each numbered clause stays whole, and the answer becomes correct. No change to the request path was needed.
Follow-up questions to expect
- "Where does query rewriting fit?" — In the online half, before retrieval. It turns "and for band 4?" into "What is the notice period for band 4 employees?" using the chat history.
- "How do you update the index without downtime?" — Upsert changed chunks by stable ID. For a full rebuild (for example, a new embedding model), build a new collection and switch an alias to it in one step.
- "Which half costs more?" — Usually the online half, because LLM calls happen per question. Indexing is paid once per document version.