Course Content
Agentic AI Patterns
9 sections · 50 lessons
What is Retrieval-Augmented Generation (RAG), and how does it support agent intelligence?
What you need to know
The pipeline
- Ingest and chunk — split documents into pieces of a few hundred tokens, keeping headings and IDs.
- Embed and index — store vectors for meaning search and text for keyword search, plus permission metadata.
- Retrieve — hybrid search (keyword such as BM25, plus vector) for the top 30 to 50 candidates, filtered by the user's permissions.
- Rerank — a reranker model picks the best 5 or so.
- Generate — answer from the chunks, with citations to chunk IDs.
Static versus agentic RAG
Static RAG
- Retrieve once, then answer
- One model call; cheap and fast
- Good for single-fact questions
- Fails when the first search misses
Agentic RAG
- Retrieval is a tool the agent calls
- Rewrites queries, follows references
- Handles multi-hop and vague questions
- More calls and latency; needs a round cap
Multi-hop means the answer needs two or more lookups where the second depends on the first: "Does my policy cover the hospital my father was admitted to?" needs the policy's network-hospital list and then the hospital's details.
Failure modes
- Retrieval misses the right chunk. Most RAG quality problems are retrieval problems. Measure recall at k on a labelled set.
- Chunking splits the answer across two pieces.
- Stale index. Yesterday's policy wording.
- Injected instructions in retrieved documents. Treat retrieved text as data.
Controls
Evaluate retrieval separately from generation, rerank, require citations and check that cited chunks exist, filter by permissions inside the query, and cap an agent at, say, 4 retrieval rounds.
A real-life example
A procurement agent answers buyers' questions about vendor contracts. A buyer asks: "Can we cancel the Vendor B laptop order without penalty if delivery is late?"
Static RAG retrieved the general cancellation clause from Vendor B's master agreement and answered "no, 10% penalty". Wrong: an amendment signed later added a late-delivery exception.
The agentic version:
- Searches "Vendor B cancellation penalty" and finds clause 14.2 of the master agreement.
- Notices clause 14.2 says "subject to amendments" and searches "Vendor B amendment cancellation".
- Finds Amendment 3, clause 2: penalty waived if delivery is more than 10 days late.
- Calls
get_order_statusand sees delivery is 12 days late. - Answers: "Yes, cancellation is penalty-free under Amendment 3, clause 2", with both citations.
Four retrieval and tool calls, about 6 seconds, instead of one wrong answer in 2 seconds.
Follow-up questions to expect
- "How do you evaluate RAG?" — Separately: retrieval with recall and precision at k on labelled queries; generation with faithfulness to the retrieved chunks and answer correctness.
- "When is static RAG enough?" — Single-fact questions over a well-chunked corpus, like "what is the leave policy for interns?". It is cheaper and more predictable.
- "Does a long context window replace RAG?" — Only for small, stable corpora. For large or permissioned data, retrieval is cheaper, faster and respects access control.