Advanced RAG

Course Content

Advanced RAG

3 sections · 38 lessons

How does indexing hypothetical questions improve retrieval?


What you need to know

The asymmetry problem

Questions and answers look different. A user asks, "Can I open an account for my 10-year-old?" The policy says, "Accounts in the name of minors may be opened with the guardian as operator, subject to the documentation in Annex 4." They share almost no words, and one is a casual question while the other is formal prose. Their embeddings end up only moderately similar, and a less relevant but question-shaped passage — say, an FAQ about adult accounts — may rank higher.

How it works

  1. Generate — for each chunk, prompt a small LLM: "Write 3–5 different questions a customer or employee might ask that this passage answers. Use plain, varied wording."
  2. Index — embed each question; metadata holds the parent chunk_id. Many teams also keep the chunk's own embedding.
  3. Search — embed the user's question, search the question vectors (and chunk vectors), map hits to chunks, and deduplicate.
  4. Generate the answer from the real chunk text — never from the generated questions.

HyDE: the query-time mirror

HyDE (Hypothetical Document Embeddings, 2022) asks an LLM to write a plausible answer to the user's question, embeds that answer, and searches with it. Answer-shaped text matches answer-shaped passages.

Hypothetical questions

  • LLM cost at index time, paid once
  • No extra latency per query
  • Index grows several times
  • Best when question styles are predictable

HyDE

  • LLM call on every query
  • Adds hundreds of ms per query
  • No index changes; easy to try
  • Can drift when the model doesn't know the domain

HyDE's risk is specific: on niche or private topics, the model's hypothetical answer may be confidently wrong, and it then retrieves documents that match the wrong answer. Hypothetical questions don't have this problem, because they are generated from real text.

Costs of hypothetical questions

  • An LLM call per chunk at ingestion, and again when the chunk changes.
  • An index several times larger (3–5 question vectors per chunk).
  • Generated questions can be too similar to each other; ask for varied wording and personas.

A real-life example

A bank's compliance assistant serves branch staff. Their questions are informal: "Can a minor open an account with just Aadhaar?", "Do we need a fresh KYC if a customer changes their address?" The policy manuals are formal and long.

The team generates four questions per chunk for 18,000 chunks with a small model — an overnight batch job. For the minor-accounts chunk, the generated questions include:

Text
Can a minor open a savings account?Who operates an account opened for a child?What documents are needed to open an account for a minor?Can a guardian open an account in a child's name?

"Can a minor open an account with just Aadhaar?" now matches the third question almost exactly, and the correct chunk ranks first. On 300 real staff questions, recall@5 improves most for short, informal questions and hardly at all for questions that quote policy terms directly — those were already matched well by the chunk vectors and by BM25.

The team also tries HyDE. It performs well on general banking topics but, on the bank's own internal procedures, the hypothetical answers invent plausible-sounding steps and pull in wrong chunks. They keep hypothetical questions and drop HyDE.

Follow-up questions to expect

  • "Do you replace the chunk embeddings with question embeddings?" — Usually no. Keep both, and treat them as extra vectors pointing at the same chunk, or fuse the two result lists with RRF.
  • "How do you keep generated questions fresh?" — Regenerate for a chunk whenever its content hash changes, as part of incremental indexing.
  • "How does this relate to multi-vector retrieval?" — It is one form of it: several representations (chunk, summary, questions) of the same content, all mapping to one parent.