Course Content
Advanced RAG
3 sections · 38 lessons
What are the pros and cons of tightly vs loosely coupled retriever-generator systems?
What you need to know
What "coupling" means here
The retriever and the generator can be connected in two ways:
Tightly coupled
- Trained together, or retrieval built into the model
- Retriever optimised for "helps the generator answer"
- Changing one part means retraining the other
- Examples: REALM, RAG (2020), RETRO, Atlas
Loosely coupled
- Separate components, joined by a text prompt
- Retriever optimised for "looks similar to the query"
- Swap the LLM or embedder independently
- Examples: almost every production RAG system today
- REALM and the original RAG paper (2020) trained the query encoder together with the reader or generator, so gradients from answer quality improved retrieval.
- RETRO built retrieval into the architecture: the model attends to retrieved neighbour chunks through special cross-attention layers.
- Atlas and later RA-DIT fine-tuned both sides so the generator learns to use retrieved text and the retriever learns to prefer passages that help.
Why tight coupling can be better
Similarity is not the same as usefulness. The passage most similar to "how do I port my number?" may be a marketing page that repeats the question; the useful one is the step-by-step guide. A jointly trained retriever learns that difference from answer quality. It also wastes fewer context tokens on passages the generator cannot use.
Why most teams stay loosely coupled
- Swappability. New frontier models arrive several times a year. With loose coupling you switch the LLM with a config change. With tight coupling you retrain.
- Hosted models. You cannot jointly train a closed API model with your retriever.
- Debuggability. You can measure retrieval recall and generation faithfulness separately, and know which half to fix.
- Cost. Joint training needs labelled data, GPUs and ML engineers; re-embedding the corpus after each retriever change is also real work.
The practical middle ground
Close the gap partially and cheaply:
- Fine-tune the embedding model on your own (query, relevant passage) pairs — from clicks, thumbs-up, or resolved tickets — using a contrastive loss.
- Fine-tune or pick a stronger reranker, which directly learns "is this passage useful for this query".
- Add query rewriting, which adapts the query to the corpus without training anything.
These keep the components separate while tuning them to your data.
A real-life example
An Indian telecom's support bot is loosely coupled: a general multilingual embedding model, a vector store, a reranker and a hosted LLM. Evaluation shows that the LLM answers well when the right article is in its context, but recall@5 is weak for Hinglish messages such as "port karna hai dusre network pe".
The team considers joint training but rejects it: they plan to upgrade the LLM within six months, and they cannot train the hosted model anyway. Instead they collect 40,000 (customer message, article that resolved the chat) pairs from support logs, clean them, and fine-tune only the embedding model. They re-embed the 12,000 articles once — a few hours of batch work — and re-run the same evaluation set. Recall for Hinglish messages improves noticeably, while nothing else in the system changes. When the LLM is upgraded later, it is a one-line configuration change.
Follow-up questions to expect
- "Where does fine-tuning the generator on retrieved context fit?" — It is a partial form of coupling: the generator learns to ignore noisy passages from your retriever. Useful, but you must redo it if the retriever changes a lot.
- "How do you get training pairs without labels?" — From implicit feedback (clicked or cited documents, resolved tickets) or synthetic questions generated from your chunks, with a human spot-check.
- "What breaks when you swap the embedding model?" — The whole index must be re-embedded, and every similarity threshold must be re-tuned. Plan it as a migration.