Retrieval-Augmented Generation (RAG)

Course Content

Retrieval-Augmented Generation (RAG)

4 sections · 8 lessons

Query Embedding Optimization


A team shipped a RAG assistant over their internal documentation and it tested beautifully. Then they instrumented it and broke the query log down by length. The picture changed completely.

Text
Query length     Share of traffic   Recall@51-3 words             34%             0.414-7 words             41%             0.728+ words              25%             0.86

A third of all traffic was being served at 41% recall. Nobody had noticed, because the evaluation set had been written by engineers, and engineers write careful eight-word questions. Real users type refund window, or why slow, or it broke again.

The instinct at this point is to blame the embedding model and go shopping for a better one. That instinct is usually wrong and always expensive. The passages were fine. The index was fine. The problem was upstream of both: the raw query, exactly as the user typed it, is often a poor search key, and there is a great deal you can do about that before the vector ever touches the index.

Four repairs for a query that retrieves badlyThe raw queryunderperformsExpansion: add likely termsRewriting:resolve 'it' and 'that'Decomposition:split compound asksHyDE: embed afabricated answerMulti-query: fanout, then fuse
Every one of these attacks the same asymmetry: a short question and a long passage sit in different parts of the embedding space even when they are about the same thing.

Why raw queries underperform

The asymmetry problem

You are comparing a 3-token query against a 400-token passage using a single similarity number. These two texts are not the same kind of object. The passage is a dense, self-contained explanation; the query is a fragment. Their vectors are computed by the same encoder but they occupy different regions of the space — queries cluster with questions, passages cluster with prose.

Many modern embedding models handle this explicitly by being trained asymmetrically, with instruction prefixes such as "query: " and "passage: ". If your model expects those prefixes and you omit them, you lose several points of recall silently, with no error and no warning.

The vocabulary gap

Users say "my money back"; the document says "reimbursement of prepaid fees". Users say "it's dead"; the runbook says "service unresponsive following OOM termination". Dense retrieval closes some of this gap, but the further your domain vocabulary sits from the model's training distribution, the wider the gap stays.

Context dependence

"Does that apply to annual plans too?" is a perfectly clear question inside a conversation and a meaningless string on its own. The words carrying the topic — refunds, contracts, the specific plan being discussed — are three turns back in the chat history and absent from the vector entirely.

Compound questions

"How do our refund terms compare between monthly and enterprise annual plans?" requires two passages that may live in two documents. Embedding the whole question produces a vector sitting between the two topics, near neither, and retrieval returns five mediocre chunks about "plans" in general.

A compound question embeds to the average of its parts, and the average of two specific things is one vague thing. Retrieval on an averaged vector is retrieval on a query nobody asked.

What we are trying to fix

GoalSymptom when unmetTechnique that addresses it
Close the vocabulary gapCorrect doc exists, never retrieved; uses different wordsQuery expansion, HyDE
Make short queries specificShort queries retrieve generic overview pagesQuery rewriting, expansion
Resolve conversational referencesFollow-up turns retrieve nonsenseQuery rewriting with history
Handle multi-part questionsAnswer covers half the questionDecomposition
Match query-space to passage-spaceUniformly mediocre scores across the boardCorrect prefixes, HyDE, fine-tuning

Technique 1 — Query expansion

Add terms to the query so it overlaps more of the vocabulary the answer might use.

Lexical expansion, no model required

Maintain a domain dictionary. This is unglamorous and extremely effective, because your domain has a fixed set of synonym pairs that a general-purpose embedding model was never taught.

Python
DOMAIN_SYNONYMS = {    "refund":   ["reimbursement", "money back", "credit note", "chargeback"],    "sso":      ["single sign-on", "saml", "oidc", "federated login"],    "slow":     ["latency", "degraded performance", "timeout", "p99"],    "cancel":   ["terminate", "churn", "non-renewal", "wind down"],}def expand_lexical(query, max_added=4):    words = query.lower().split()    added = []    for w in words:        for syn in DOMAIN_SYNONYMS.get(w.strip("?.,"), []):            if syn not in query.lower() and len(added) < max_added:                added.append(syn)    return query if not added else f"{query} ({', '.join(added)})"expand_lexical("refund window")# 'refund window (reimbursement, money back, credit note, chargeback)'

The caveat is real: every added term shifts the query vector. Add eight synonyms and the vector drifts towards the centroid of a synonym cloud rather than the user's actual question. Cap the additions, and keep the original words at the front where they carry most weight.

LLM-based expansion

Ask a small, cheap model to enrich the query with likely terminology.

Python
import os# small, cheap model for query-side work; the name lives in config.# No token cap: on reasoning models, reasoning tokens count against# max_completion_tokens, and a small cap can return an empty reply.FAST_MODEL = os.getenv("FAST_MODEL", "gpt-6-luna")EXPAND = """Rewrite this search query to include terminology that wouldappear in an internal policy document answering it. Keep it under 30 words.Output only the rewritten query.Query: {q}"""def expand_llm(client, q):    r = client.chat.completions.create(        model=FAST_MODEL,        messages=[{"role": "user", "content": EXPAND.format(q=q)}])    return r.choices[0].message.content.strip()# "refund window"# -> "refund window: eligibility period for reimbursement of prepaid fees#     under enterprise and monthly subscription contracts"

That expansion costs one small-model call — a few hundred milliseconds and a fraction of a cent — and on a short-query bucket like the one in the table at the top it is usually the best return available. Measure the lift on your own labelled queries rather than assuming it.

Technique 2 — Query rewriting

Expansion adds terms; rewriting replaces the query with a better-formed standalone question. It is the fix for conversational context and for garbled input.

Python
REWRITE = """Given the conversation, rewrite the final user message as astandalone search query. Resolve all pronouns and references explicitly.Correct obvious spelling errors. Output only the query.Conversation:{history}Final message: {q}"""def rewrite(client, q, history):    hist = "\n".join(f"{m['role']}: {m['content']}" for m in history[-6:])    r = client.chat.completions.create(        model=FAST_MODEL,        messages=[{"role": "user",                   "content": REWRITE.format(history=hist, q=q)}])    return r.choices[0].message.content.strip()
Conversation so farUser typesRewritten query
Discussion of enterprise annual refund terms"what about monthly?""What is the refund window for monthly subscription plans?"
Troubleshooting a failing SAML login"it broke again""SAML single sign-on login failure troubleshooting"
None"kubrnetes pod crashloop""Kubernetes pod CrashLoopBackOff diagnosis"

Rewriting is close to mandatory for any multi-turn interface. Without it, every follow-up question retrieves on a fragment, and users experience the assistant as having no memory — which, from retrieval's point of view, it does not.

Technique 3 — Query decomposition

Split a compound question into independent sub-questions, retrieve for each, merge the results.

Python
DECOMPOSE = """Split the question into the minimum set of independentsub-questions needed to answer it. If it is already a single question,return it unchanged. One per line, no numbering.Question: {q}"""def decompose(client, q):    r = client.chat.completions.create(        model=FAST_MODEL,        messages=[{"role": "user", "content": DECOMPOSE.format(q=q)}])    return [line.strip() for line in r.choices[0].message.content.splitlines()            if line.strip()]def retrieve_decomposed(client, retriever, q, k_each=3, k_final=6):    subs = decompose(client, q)    pooled = {}    for sub in subs:        for hit in retriever.search(sub, k=k_each):            # keep the best score any sub-question achieved for this chunk            prev = pooled.get(hit["id"])            if prev is None or hit["score"] > prev["score"]:                pooled[hit["id"]] = hit    return sorted(pooled.values(), key=lambda h: -h["score"])[:k_final]

Applied to "How do our refund terms compare between monthly and enterprise annual plans?", decomposition yields two clean queries — one about monthly refund terms, one about enterprise annual refund terms — each of which retrieves its own precise chunk. The merged context now contains both facts, and the model can actually perform the comparison it was asked for.

The cost is one extra model call plus N retrievals. Do not apply it unconditionally; a cheap heuristic (does the query contain "and", "compare", "versus", "both", or more than one question mark?) routes only the queries that need it.

Technique 4 — HyDE: search with a fake answer

HyDE — Hypothetical Document Embeddings — is the most counter-intuitive technique here and often the most effective. The idea: do not embed the question. Ask the model to invent an answer, then embed that.

Why on earth would that help, given the invented answer may be factually wrong? Because you are no longer comparing a question to a passage. You are comparing a passage-shaped text to a passage. The hypothetical answer has the length, register, structure and vocabulary of a real document — so it lands in the same region of embedding space as real documents. Its factual errors do not matter, because it is thrown away immediately after being embedded. It never reaches the user.

HyDE works because retrieval quality depends on the shape of the text you embed, not its truth. A wrong answer that looks like a document beats a right question that does not.

Python
HYDE = """Write a short factual passage (about 80 words) that would appear ina company policy document and would answer this question. Write it asdocumentation, not as an answer to a person. Invent plausible specifics.Question: {q}"""def hyde_search(client, retriever, q, k=5):    r = client.chat.completions.create(        model=FAST_MODEL,        messages=[{"role": "user", "content": HYDE.format(q=q)}])    hypothetical = r.choices[0].message.content.strip()    # embed the fake passage, not the question    return retriever.search(hypothetical, k=k)

An illustrative before-and-after on the query "refund window", against a corpus of policy documents:

What is embeddedCosine with the correct chunkRank of correct chunk
Raw query: "refund window"0.617
LLM-expanded query0.743
HyDE hypothetical passage0.831

HyDE's costs are honest ones: an extra model call (roughly 300–600 ms and a fraction of a cent), and a failure mode where the hypothetical drifts into a topic your corpus does not contain, dragging retrieval with it. On very obscure queries — where the model has no idea what a plausible answer looks like — HyDE can be worse than the raw query. A robust configuration retrieves with both the raw query and the hypothetical and fuses the two result lists.

Multi-query retrieval, and why ColBERT is a different thing

These two get confused constantly, and they operate at completely different levels.

Multi-query retrieval generates several full paraphrases of the question, embeds each into its own vector, runs several searches, and fuses the ranked lists. It is a query-side technique, works with any off-the-shelf index, and costs one model call plus N searches.

Python
MULTI = """Generate 3 different phrasings of this question, each usingdifferent vocabulary. One per line, no numbering.Question: {q}"""def multi_query(client, retriever, q, k=5):    r = client.chat.completions.create(        model=FAST_MODEL,        messages=[{"role": "user", "content": MULTI.format(q=q)}])    variants = [q] + [l.strip() for l in                      r.choices[0].message.content.splitlines() if l.strip()]    rankings = [[h["id"] for h in retriever.search(v, k=20)] for v in variants]    # reciprocal rank fusion across the variant rankings    scores = {}    for ranking in rankings:        for rank, doc_id in enumerate(ranking, start=1):            scores[doc_id] = scores.get(doc_id, 0.0) + 1.0 / (60 + rank)    return sorted(scores, key=lambda d: -scores[d])[:k]

ColBERT is an index architecture, not a query trick. Instead of one vector per text, it stores one vector per token, and scores a query-document pair by "late interaction" — for each query token, find its best-matching document token, and sum those maxima:

S(q,d)=∑i∈q max⁡j∈d Eqi⋅EdjS(q,d)=\sum_{i \in q}\ \max_{j \in d}\ E_{q_i}\cdot E_{d_j}

This retains far more detail than squashing a passage into a single vector, and it is genuinely more accurate. It also multiplies your storage by the number of tokens per chunk. A 300-token chunk that took one 768-dimensional vector now takes 300 of them — a hundredfold-plus increase before compression. ColBERT implementations use aggressive quantisation to claw that back, but the operational weight is real.

Multi-queryColBERT
LevelQuery preprocessingIndex and scoring architecture
Vectors per chunk1One per token
Extra model call at query timeYesNo
Works with existing FAISS/pgvector indexYesNo, needs a ColBERT-aware store
Added latency~400 ms (generation) + N searchesHigher per-search cost, no generation
Storage impactNoneLarge

Choosing the embedding model

Model choice is the one decision that is expensive to reverse: changing it means re-embedding your entire corpus, because vectors from two different models are not comparable.

ModelDimsMax tokensWhere it fits
all-MiniLM-L6-v2384256CPU-only, high volume, small corpora; weakest on jargon
BAAI/bge-base-en-v1.5768512Strong open default; needs a query instruction prefix
intfloat/e5-large-v21024512Higher accuracy, heavier; needs "query:"/"passage:" prefixes
text-embedding-3-small15368192Hosted, no infra, long inputs, dimensions truncatable
text-embedding-3-large30728192Top hosted quality; four times the storage of 768-dim

Two things people miss. First, the max-token limit is a silent truncator. Feed a 700-token chunk to a 512-token model and the last 190 tokens are discarded without complaint — you have indexed a document you think contains a fact that is not in the index at all. Second, some models support Matryoshka truncation: you can slice a 1536-dimensional vector down to 512 and keep most of the quality, cutting storage by two-thirds. Check the model card before you pay for dimensions you do not need.

Choosing by constraint

Binding constraintChooseWhy
Data cannot leave your networkOpen model, self-hostedHosted APIs are off the table entirely
Corpus above 10M chunks384 or truncated 512 dimsMemory dominates cost; 10M x 1536 x 4 bytes is 61 GB
Heavy domain jargonOpen model + fine-tuningOnly path to teaching the model your vocabulary
Prototype, deadline this weekHosted APINo GPU, no serving, no ops
Multilingual usersMultilingual-E5 or similarEnglish-only models collapse across languages
Chunks longer than 512 tokensLong-context embedding modelOtherwise you index silently truncated text

Fine-tuning for your domain

When a general model genuinely does not understand your vocabulary, fine-tuning is the fix. You need query-passage pairs, and the good news is you can generate them: for each chunk, ask a model to write three questions that chunk answers. A few thousand pairs is enough to move the needle.

Python
from sentence_transformers import SentenceTransformer, InputExample, lossesfrom torch.utils.data import DataLoadermodel = SentenceTransformer("BAAI/bge-small-en-v1.5")examples = [InputExample(texts=[q, passage]) for q, passage in pairs]loader = DataLoader(examples, shuffle=True, batch_size=32)# every other passage in the batch acts as an in-batch negativeloss = losses.MultipleNegativesRankingLoss(model)model.fit(train_objectives=[(loader, loss)], epochs=2, warmup_steps=100)model.save("bge-small-ourdomain")

The loss function is the clever part: it needs no explicitly labelled negatives, because within a batch of 32 pairs, the other 31 passages are treated as negatives for each query. Larger batches therefore give a harder and more informative training signal, which is why batch size matters more here than in most fine-tuning.

Measuring whether any of this helped

None of these techniques is free, and several can make things worse on the wrong corpus. You need a labelled set — 50 real queries, each paired with the id of the chunk that answers it — and two numbers.

Recall@k: the fraction of queries whose correct chunk appears in the top k. MRR@k: the mean of 1/rank of the first correct chunk, counting zero when it is absent.

Six queries, ranks of the correct chunk before and after adding LLM expansion plus HyDE fusion:

QueryRank, rawRank, optimised
"refund window"31
"why slow"not in top 104
"how to enable sso"11
"cancel mid-term"72
"data retention"21
"it broke again"not in top 106

Recall@5 before: three of six queries had the answer inside the top five, so 3/6 = 0.500. After: five of six, so 5/6 = 0.833.

MRR@10 before: 1/3 + 0 + 1 + 1/7 + 1/2 + 0 = 0.3333 + 1 + 0.1429 + 0.5 = 1.9762, divided by 6 = 0.329. After: 1 + 1/4 + 1 + 1/2 + 1 + 1/6 = 3.9167, divided by 6 = 0.653.

Recall rose 33 points and MRR doubled. That is a genuine, defensible improvement — and note it required no change to the embedding model, the chunking, or the index.

Python
def evaluate(retriever, labelled, k=5, mrr_at=10):    hits, rr = 0, 0.0    for query, gold_chunk_id in labelled:        results = retriever.search(query, k=mrr_at)        ids = [r["id"] for r in results]        if gold_chunk_id in ids[:k]:            hits += 1        if gold_chunk_id in ids:            rr += 1.0 / (ids.index(gold_chunk_id) + 1)    n = len(labelled)    return {"recall_at_k": hits / n, "mrr": rr / n, "n": n}print(evaluate(base_retriever, labelled))       # {'recall_at_k': 0.5,   'mrr': 0.329}print(evaluate(optimised_retriever, labelled))  # {'recall_at_k': 0.833, 'mrr': 0.653}

Where people get this wrong

MistakeWhat happensCorrect approach
Different prefix (or model) for queries and passagesQuality drops several points, no error raisedOne encoder, correct prefixes per the model card, applied consistently
Applying every technique at onceLatency triples; you cannot tell what helpedAdd one at a time, measure Recall@k after each
Expanding with 15 synonymsVector drifts to a synonym centroid; precision fallsCap at 3–5 additions; keep original terms first
Running HyDE on every query+500 ms on queries that were already fineRoute by heuristic: short or low-scoring queries only
Evaluating on engineer-written questionsMetrics look great, production does notSample the real query log, including the ugly ones
Chunks longer than the model's token limitText silently truncated; facts unreachableAssert chunk token length against the model limit at index time
Rewriting with the full chat historyOld topics leak in and pollute the queryUse the last 3–6 turns only

What this means when you build one

Start by looking at your query log the way that team did — bucket by length, or by whether the query is a fragment or a sentence, and compute recall per bucket. Aggregate metrics hide the failure. A system at 0.72 overall recall may be at 0.41 on a third of its traffic, and that third is where your angry users are.

Then apply techniques selectively rather than universally, because every one of them costs latency. A pragmatic router: if the query is under four words or the top retrieval score comes back below your calibrated threshold, spend the extra model call on rewriting or HyDE; otherwise go straight to the index. Most queries need nothing, and paying 500 ms on all of them to fix a third of them is a bad trade you can avoid with one if.

Above all, treat query optimisation as the cheap lever it is. Changing the embedding model means re-embedding the corpus, revalidating everything, and re-tuning thresholds. Adding a rewriting step is twenty lines of code and can be switched off in a deploy. Exhaust the cheap lever before you reach for the expensive one — and measure both against the same fifty labelled queries so you know which one actually earned its place.