RAG Systems

Course Content

RAG Systems

12 sections · 66 lessons

How is retrieved context injected into the prompt?


What the model actually receivesSystem: rules, citation format, exact refusal linedoc 1, spec_sheet — strongest chunk firstdoc 2 to doc 5 — labelled, about 250 tokens eachQuestion last, right before the answer starts
The model never sees the vector store, only this text — labels and order are what make citations and checks possible.

What you need to know

"Augmented" in RAG means exactly this step. Retrieval returns a list of Document objects. The model can read only a prompt. So something has to turn the list into text and place that text in the right spot.

The four parts of a RAG prompt

  1. Rules (system message): who the assistant is, "use only the documents", how to cite, what to say when the answer is missing.
  2. Documents: the retrieved chunks, each clearly separated and labelled.
  3. Question: the user's words, placed after the documents.
  4. Output format (optional): for example, "answer in at most 3 sentences, then list sources".

Putting the long documents first and the question at the end is the layout Anthropic recommends for long inputs, and OpenAI's long-context advice is to repeat key instructions after the documents too. The model reads the question right before it starts answering, with the evidence above it.

The code

Here is the injection step as a LangChain LCEL chain. I ran it with a fake retriever and printed the rendered prompt.

Python
from langchain_core.output_parsers import StrOutputParserfrom langchain_core.prompts import ChatPromptTemplatefrom langchain_core.runnables import RunnablePassthroughSYSTEM = """You answer questions about products sold on our store.Use only the documents in <documents>. Cite them like [2].If they do not contain the answer, reply: I could not find that in the product details."""HUMAN = "<documents>\n{context}\n</documents>\n\nQuestion: {question}"prompt = ChatPromptTemplate.from_messages([("system", SYSTEM), ("human", HUMAN)])def format_docs(docs):    return "\n".join(        f'<doc id="{i}" source="{d.metadata["source"]}">\n{d.page_content}\n</doc>'        for i, d in enumerate(docs, start=1))chain = ({"context": retriever | format_docs, "question": RunnablePassthrough()}         | prompt | llm | StrOutputParser())

It assumes you already have a retriever and a chat model llm. The dictionary at the start runs two things on the same input: the retriever gets the question and its documents are formatted, while RunnablePassthrough() passes the question through unchanged. Both fill the template, and the chain sends the result to the model.

Why each detail matters

  • Delimiters (<doc> tags, or --- lines) stop two chunks from blending into one, and make it harder for text inside a chunk to pose as instructions.
  • Numbers and sources let the model cite [2], and let your code check that [2] really exists.
  • Metadata in the tag, such as effective_date="2026-04-01", lets the model prefer the newer source when two disagree.
  • A token budget. Five chunks of about 250 tokens each is roughly 1,250 tokens. That is enough for most factual questions and keeps cost and delay low.

A real-life example

An e-commerce site answers product questions under each listing. A shopper asks, "Can I charge it wirelessly?" Retrieval returns a spec-sheet chunk and a seller FAQ chunk. The rendered prompt looks like this:

Text
<documents><doc id="1" source="spec_sheet">Battery: 5,000 mAh. Charger in box: 33 W.</doc><doc id="2" source="seller_faq">Q: Does it support wireless charging? A: No.</doc></documents>Question: Can I charge it wirelessly?

The model answers "No, this phone does not support wireless charging [2]." The page shows "Source: seller FAQ" under the answer. In an earlier version the team joined chunks with a single newline and no labels. The model then mixed the 33 W charger detail into a wrong claim of "33 W wireless charging". The labels fixed that without changing the model.

Follow-up questions to expect

  • "Where should the instructions go, before or after the documents?" — Core rules go in the system message. With very long contexts, repeating the key instruction and the question after the documents helps, because the end of the prompt is read last.
  • "How many chunks do you inject?" — Retrieve 20 to 50 candidates, rerank, and inject the top 3 to 6. Then tune on an evaluation set, not by feel.
  • "Does the order of chunks matter?" — Yes. Put the strongest chunk first, or first and last, because models use information at the edges of a long context more reliably than information in the middle.