Course Content
Retrieval-Augmented Generation (RAG)
4 sections · 8 lessons
Query Embedding Optimization
A team shipped a RAG assistant over their internal documentation and it tested beautifully. Then they instrumented it and broke the query log down by length. The picture changed completely.
Query length Share of traffic Recall@51-3 words 34% 0.414-7 words 41% 0.728+ words 25% 0.86A third of all traffic was being served at 41% recall. Nobody had noticed, because the evaluation set had been written by engineers, and engineers write careful eight-word questions. Real users type refund window, or why slow, or it broke again.
The instinct at this point is to blame the embedding model and go shopping for a better one. That instinct is usually wrong and always expensive. The passages were fine. The index was fine. The problem was upstream of both: the raw query, exactly as the user typed it, is often a poor search key, and there is a great deal you can do about that before the vector ever touches the index.
Why raw queries underperform
The asymmetry problem
You are comparing a 3-token query against a 400-token passage using a single similarity number. These two texts are not the same kind of object. The passage is a dense, self-contained explanation; the query is a fragment. Their vectors are computed by the same encoder but they occupy different regions of the space — queries cluster with questions, passages cluster with prose.
Many modern embedding models handle this explicitly by being trained asymmetrically, with instruction prefixes such as "query: " and "passage: ". If your model expects those prefixes and you omit them, you lose several points of recall silently, with no error and no warning.
The vocabulary gap
Users say "my money back"; the document says "reimbursement of prepaid fees". Users say "it's dead"; the runbook says "service unresponsive following OOM termination". Dense retrieval closes some of this gap, but the further your domain vocabulary sits from the model's training distribution, the wider the gap stays.
Context dependence
"Does that apply to annual plans too?" is a perfectly clear question inside a conversation and a meaningless string on its own. The words carrying the topic — refunds, contracts, the specific plan being discussed — are three turns back in the chat history and absent from the vector entirely.
Compound questions
"How do our refund terms compare between monthly and enterprise annual plans?" requires two passages that may live in two documents. Embedding the whole question produces a vector sitting between the two topics, near neither, and retrieval returns five mediocre chunks about "plans" in general.
A compound question embeds to the average of its parts, and the average of two specific things is one vague thing. Retrieval on an averaged vector is retrieval on a query nobody asked.
What we are trying to fix
| Goal | Symptom when unmet | Technique that addresses it |
|---|---|---|
| Close the vocabulary gap | Correct doc exists, never retrieved; uses different words | Query expansion, HyDE |
| Make short queries specific | Short queries retrieve generic overview pages | Query rewriting, expansion |
| Resolve conversational references | Follow-up turns retrieve nonsense | Query rewriting with history |
| Handle multi-part questions | Answer covers half the question | Decomposition |
| Match query-space to passage-space | Uniformly mediocre scores across the board | Correct prefixes, HyDE, fine-tuning |
Technique 1 — Query expansion
Add terms to the query so it overlaps more of the vocabulary the answer might use.
Lexical expansion, no model required
Maintain a domain dictionary. This is unglamorous and extremely effective, because your domain has a fixed set of synonym pairs that a general-purpose embedding model was never taught.
1DOMAIN_SYNONYMS = {2 "refund": ["reimbursement", "money back", "credit note", "chargeback"],3 "sso": ["single sign-on", "saml", "oidc", "federated login"],4 "slow": ["latency", "degraded performance", "timeout", "p99"],5 "cancel": ["terminate", "churn", "non-renewal", "wind down"],6}78def expand_lexical(query, max_added=4):9 words = query.lower().split()10 added = []11 for w in words:12 for syn in DOMAIN_SYNONYMS.get(w.strip("?.,"), []):13 if syn not in query.lower() and len(added) < max_added:14 added.append(syn)15 return query if not added else f"{query} ({', '.join(added)})"1617expand_lexical("refund window")18# 'refund window (reimbursement, money back, credit note, chargeback)'The caveat is real: every added term shifts the query vector. Add eight synonyms and the vector drifts towards the centroid of a synonym cloud rather than the user's actual question. Cap the additions, and keep the original words at the front where they carry most weight.
LLM-based expansion
Ask a small, cheap model to enrich the query with likely terminology.
1import os23# small, cheap model for query-side work; the name lives in config.4# No token cap: on reasoning models, reasoning tokens count against5# max_completion_tokens, and a small cap can return an empty reply.6FAST_MODEL = os.getenv("FAST_MODEL", "gpt-6-luna")78EXPAND = """Rewrite this search query to include terminology that would9appear in an internal policy document answering it. Keep it under 30 words.10Output only the rewritten query.1112Query: {q}"""1314def expand_llm(client, q):15 r = client.chat.completions.create(16 model=FAST_MODEL,17 messages=[{"role": "user", "content": EXPAND.format(q=q)}])18 return r.choices[0].message.content.strip()1920# "refund window"21# -> "refund window: eligibility period for reimbursement of prepaid fees22# under enterprise and monthly subscription contracts"That expansion costs one small-model call — a few hundred milliseconds and a fraction of a cent — and on a short-query bucket like the one in the table at the top it is usually the best return available. Measure the lift on your own labelled queries rather than assuming it.
Technique 2 — Query rewriting
Expansion adds terms; rewriting replaces the query with a better-formed standalone question. It is the fix for conversational context and for garbled input.
1REWRITE = """Given the conversation, rewrite the final user message as a2standalone search query. Resolve all pronouns and references explicitly.3Correct obvious spelling errors. Output only the query.45Conversation:6{history}78Final message: {q}"""910def rewrite(client, q, history):11 hist = "\n".join(f"{m['role']}: {m['content']}" for m in history[-6:])12 r = client.chat.completions.create(13 model=FAST_MODEL,14 messages=[{"role": "user",15 "content": REWRITE.format(history=hist, q=q)}])16 return r.choices[0].message.content.strip()| Conversation so far | User types | Rewritten query |
|---|---|---|
| Discussion of enterprise annual refund terms | "what about monthly?" | "What is the refund window for monthly subscription plans?" |
| Troubleshooting a failing SAML login | "it broke again" | "SAML single sign-on login failure troubleshooting" |
| None | "kubrnetes pod crashloop" | "Kubernetes pod CrashLoopBackOff diagnosis" |
Rewriting is close to mandatory for any multi-turn interface. Without it, every follow-up question retrieves on a fragment, and users experience the assistant as having no memory — which, from retrieval's point of view, it does not.
Technique 3 — Query decomposition
Split a compound question into independent sub-questions, retrieve for each, merge the results.
1DECOMPOSE = """Split the question into the minimum set of independent2sub-questions needed to answer it. If it is already a single question,3return it unchanged. One per line, no numbering.45Question: {q}"""67def decompose(client, q):8 r = client.chat.completions.create(9 model=FAST_MODEL,10 messages=[{"role": "user", "content": DECOMPOSE.format(q=q)}])11 return [line.strip() for line in r.choices[0].message.content.splitlines()12 if line.strip()]1314def retrieve_decomposed(client, retriever, q, k_each=3, k_final=6):15 subs = decompose(client, q)16 pooled = {}17 for sub in subs:18 for hit in retriever.search(sub, k=k_each):19 # keep the best score any sub-question achieved for this chunk20 prev = pooled.get(hit["id"])21 if prev is None or hit["score"] > prev["score"]:22 pooled[hit["id"]] = hit23 return sorted(pooled.values(), key=lambda h: -h["score"])[:k_final]Applied to "How do our refund terms compare between monthly and enterprise annual plans?", decomposition yields two clean queries — one about monthly refund terms, one about enterprise annual refund terms — each of which retrieves its own precise chunk. The merged context now contains both facts, and the model can actually perform the comparison it was asked for.
The cost is one extra model call plus N retrievals. Do not apply it unconditionally; a cheap heuristic (does the query contain "and", "compare", "versus", "both", or more than one question mark?) routes only the queries that need it.
Technique 4 — HyDE: search with a fake answer
HyDE — Hypothetical Document Embeddings — is the most counter-intuitive technique here and often the most effective. The idea: do not embed the question. Ask the model to invent an answer, then embed that.
Why on earth would that help, given the invented answer may be factually wrong? Because you are no longer comparing a question to a passage. You are comparing a passage-shaped text to a passage. The hypothetical answer has the length, register, structure and vocabulary of a real document — so it lands in the same region of embedding space as real documents. Its factual errors do not matter, because it is thrown away immediately after being embedded. It never reaches the user.
HyDE works because retrieval quality depends on the shape of the text you embed, not its truth. A wrong answer that looks like a document beats a right question that does not.
1HYDE = """Write a short factual passage (about 80 words) that would appear in2a company policy document and would answer this question. Write it as3documentation, not as an answer to a person. Invent plausible specifics.45Question: {q}"""67def hyde_search(client, retriever, q, k=5):8 r = client.chat.completions.create(9 model=FAST_MODEL,10 messages=[{"role": "user", "content": HYDE.format(q=q)}])11 hypothetical = r.choices[0].message.content.strip()12 # embed the fake passage, not the question13 return retriever.search(hypothetical, k=k)An illustrative before-and-after on the query "refund window", against a corpus of policy documents:
| What is embedded | Cosine with the correct chunk | Rank of correct chunk |
|---|---|---|
| Raw query: "refund window" | 0.61 | 7 |
| LLM-expanded query | 0.74 | 3 |
| HyDE hypothetical passage | 0.83 | 1 |
HyDE's costs are honest ones: an extra model call (roughly 300–600 ms and a fraction of a cent), and a failure mode where the hypothetical drifts into a topic your corpus does not contain, dragging retrieval with it. On very obscure queries — where the model has no idea what a plausible answer looks like — HyDE can be worse than the raw query. A robust configuration retrieves with both the raw query and the hypothetical and fuses the two result lists.
Multi-query retrieval, and why ColBERT is a different thing
These two get confused constantly, and they operate at completely different levels.
Multi-query retrieval generates several full paraphrases of the question, embeds each into its own vector, runs several searches, and fuses the ranked lists. It is a query-side technique, works with any off-the-shelf index, and costs one model call plus N searches.
1MULTI = """Generate 3 different phrasings of this question, each using2different vocabulary. One per line, no numbering.34Question: {q}"""56def multi_query(client, retriever, q, k=5):7 r = client.chat.completions.create(8 model=FAST_MODEL,9 messages=[{"role": "user", "content": MULTI.format(q=q)}])10 variants = [q] + [l.strip() for l in11 r.choices[0].message.content.splitlines() if l.strip()]12 rankings = [[h["id"] for h in retriever.search(v, k=20)] for v in variants]13 # reciprocal rank fusion across the variant rankings14 scores = {}15 for ranking in rankings:16 for rank, doc_id in enumerate(ranking, start=1):17 scores[doc_id] = scores.get(doc_id, 0.0) + 1.0 / (60 + rank)18 return sorted(scores, key=lambda d: -scores[d])[:k]ColBERT is an index architecture, not a query trick. Instead of one vector per text, it stores one vector per token, and scores a query-document pair by "late interaction" — for each query token, find its best-matching document token, and sum those maxima:
This retains far more detail than squashing a passage into a single vector, and it is genuinely more accurate. It also multiplies your storage by the number of tokens per chunk. A 300-token chunk that took one 768-dimensional vector now takes 300 of them — a hundredfold-plus increase before compression. ColBERT implementations use aggressive quantisation to claw that back, but the operational weight is real.
| Multi-query | ColBERT | |
|---|---|---|
| Level | Query preprocessing | Index and scoring architecture |
| Vectors per chunk | 1 | One per token |
| Extra model call at query time | Yes | No |
| Works with existing FAISS/pgvector index | Yes | No, needs a ColBERT-aware store |
| Added latency | ~400 ms (generation) + N searches | Higher per-search cost, no generation |
| Storage impact | None | Large |
Choosing the embedding model
Model choice is the one decision that is expensive to reverse: changing it means re-embedding your entire corpus, because vectors from two different models are not comparable.
| Model | Dims | Max tokens | Where it fits |
|---|---|---|---|
all-MiniLM-L6-v2 | 384 | 256 | CPU-only, high volume, small corpora; weakest on jargon |
BAAI/bge-base-en-v1.5 | 768 | 512 | Strong open default; needs a query instruction prefix |
intfloat/e5-large-v2 | 1024 | 512 | Higher accuracy, heavier; needs "query:"/"passage:" prefixes |
text-embedding-3-small | 1536 | 8192 | Hosted, no infra, long inputs, dimensions truncatable |
text-embedding-3-large | 3072 | 8192 | Top hosted quality; four times the storage of 768-dim |
Two things people miss. First, the max-token limit is a silent truncator. Feed a 700-token chunk to a 512-token model and the last 190 tokens are discarded without complaint — you have indexed a document you think contains a fact that is not in the index at all. Second, some models support Matryoshka truncation: you can slice a 1536-dimensional vector down to 512 and keep most of the quality, cutting storage by two-thirds. Check the model card before you pay for dimensions you do not need.
Choosing by constraint
| Binding constraint | Choose | Why |
|---|---|---|
| Data cannot leave your network | Open model, self-hosted | Hosted APIs are off the table entirely |
| Corpus above 10M chunks | 384 or truncated 512 dims | Memory dominates cost; 10M x 1536 x 4 bytes is 61 GB |
| Heavy domain jargon | Open model + fine-tuning | Only path to teaching the model your vocabulary |
| Prototype, deadline this week | Hosted API | No GPU, no serving, no ops |
| Multilingual users | Multilingual-E5 or similar | English-only models collapse across languages |
| Chunks longer than 512 tokens | Long-context embedding model | Otherwise you index silently truncated text |
Fine-tuning for your domain
When a general model genuinely does not understand your vocabulary, fine-tuning is the fix. You need query-passage pairs, and the good news is you can generate them: for each chunk, ask a model to write three questions that chunk answers. A few thousand pairs is enough to move the needle.
1from sentence_transformers import SentenceTransformer, InputExample, losses2from torch.utils.data import DataLoader34model = SentenceTransformer("BAAI/bge-small-en-v1.5")5examples = [InputExample(texts=[q, passage]) for q, passage in pairs]6loader = DataLoader(examples, shuffle=True, batch_size=32)78# every other passage in the batch acts as an in-batch negative9loss = losses.MultipleNegativesRankingLoss(model)10model.fit(train_objectives=[(loader, loss)], epochs=2, warmup_steps=100)11model.save("bge-small-ourdomain")The loss function is the clever part: it needs no explicitly labelled negatives, because within a batch of 32 pairs, the other 31 passages are treated as negatives for each query. Larger batches therefore give a harder and more informative training signal, which is why batch size matters more here than in most fine-tuning.
Measuring whether any of this helped
None of these techniques is free, and several can make things worse on the wrong corpus. You need a labelled set — 50 real queries, each paired with the id of the chunk that answers it — and two numbers.
Recall@k: the fraction of queries whose correct chunk appears in the top k. MRR@k: the mean of 1/rank of the first correct chunk, counting zero when it is absent.
Six queries, ranks of the correct chunk before and after adding LLM expansion plus HyDE fusion:
| Query | Rank, raw | Rank, optimised |
|---|---|---|
| "refund window" | 3 | 1 |
| "why slow" | not in top 10 | 4 |
| "how to enable sso" | 1 | 1 |
| "cancel mid-term" | 7 | 2 |
| "data retention" | 2 | 1 |
| "it broke again" | not in top 10 | 6 |
Recall@5 before: three of six queries had the answer inside the top five, so 3/6 = 0.500. After: five of six, so 5/6 = 0.833.
MRR@10 before: 1/3 + 0 + 1 + 1/7 + 1/2 + 0 = 0.3333 + 1 + 0.1429 + 0.5 = 1.9762, divided by 6 = 0.329. After: 1 + 1/4 + 1 + 1/2 + 1 + 1/6 = 3.9167, divided by 6 = 0.653.
Recall rose 33 points and MRR doubled. That is a genuine, defensible improvement — and note it required no change to the embedding model, the chunking, or the index.
1def evaluate(retriever, labelled, k=5, mrr_at=10):2 hits, rr = 0, 0.03 for query, gold_chunk_id in labelled:4 results = retriever.search(query, k=mrr_at)5 ids = [r["id"] for r in results]6 if gold_chunk_id in ids[:k]:7 hits += 18 if gold_chunk_id in ids:9 rr += 1.0 / (ids.index(gold_chunk_id) + 1)10 n = len(labelled)11 return {"recall_at_k": hits / n, "mrr": rr / n, "n": n}1213print(evaluate(base_retriever, labelled)) # {'recall_at_k': 0.5, 'mrr': 0.329}14print(evaluate(optimised_retriever, labelled)) # {'recall_at_k': 0.833, 'mrr': 0.653}Where people get this wrong
| Mistake | What happens | Correct approach |
|---|---|---|
| Different prefix (or model) for queries and passages | Quality drops several points, no error raised | One encoder, correct prefixes per the model card, applied consistently |
| Applying every technique at once | Latency triples; you cannot tell what helped | Add one at a time, measure Recall@k after each |
| Expanding with 15 synonyms | Vector drifts to a synonym centroid; precision falls | Cap at 3–5 additions; keep original terms first |
| Running HyDE on every query | +500 ms on queries that were already fine | Route by heuristic: short or low-scoring queries only |
| Evaluating on engineer-written questions | Metrics look great, production does not | Sample the real query log, including the ugly ones |
| Chunks longer than the model's token limit | Text silently truncated; facts unreachable | Assert chunk token length against the model limit at index time |
| Rewriting with the full chat history | Old topics leak in and pollute the query | Use the last 3–6 turns only |
What this means when you build one
Start by looking at your query log the way that team did — bucket by length, or by whether the query is a fragment or a sentence, and compute recall per bucket. Aggregate metrics hide the failure. A system at 0.72 overall recall may be at 0.41 on a third of its traffic, and that third is where your angry users are.
Then apply techniques selectively rather than universally, because every one of them costs latency. A pragmatic router: if the query is under four words or the top retrieval score comes back below your calibrated threshold, spend the extra model call on rewriting or HyDE; otherwise go straight to the index. Most queries need nothing, and paying 500 ms on all of them to fix a third of them is a bad trade you can avoid with one if.
Above all, treat query optimisation as the cheap lever it is. Changing the embedding model means re-embedding the corpus, revalidating everything, and re-tuning thresholds. Adding a rewriting step is twenty lines of code and can be switched off in a deploy. Exhaust the cheap lever before you reach for the expensive one — and measure both against the same fifty labelled queries so you know which one actually earned its place.