RAG Systems

Course Content

RAG Systems

12 sections · 66 lessons

What is RAG with feedback loops?


What you need to know

Two kinds of loop

  • Within a request: automatic, bounded, takes milliseconds to seconds: the corrective and self-checking loops from the earlier lessons.
  • Across requests: slower, involves people, takes days or weeks, but decides whether the system is better in six months than on launch day.

Signals to collect

ExplicitImplicit
Thumbs up / downUser rephrases the same question within a minute
"This didn't answer my question"User opens the cited source (answer may have been incomplete)
Escalation to a human agentUser copies the answer (usually a good sign)
Written commentUser leaves right after the answer

Implicit signals are noisy one by one, but useful in bulk.

The loop

  1. Capture — store feedback with the trace id of the answer.
  2. Sample — negative feedback plus a random sample of positive, so you see both.
  3. Triage — label the first failing stage: missing content, retrieval, ranking, generation, stale data.
  4. Fix at the source — add or update documents, change chunking, add metadata or synonyms, adjust the prompt.
  5. Lock it in — add the case to the golden set so it can never silently break again.
  6. Learn from volume — use logged queries with good and bad chunks to fine-tune the reranker or embedding model.

Worked numbers

An HR assistant handles 50,000 questions a week, and 2% get a thumbs-down: 1,000. The team reads a sample of 100 each week. Typical split for a mature system: many "content missing" (the policy does not exist or is vague), then retrieval misses, then generation problems. The content group goes to HR as a list of documents to write or fix. That is often the highest-value outcome of the whole loop: the RAG system shows the organisation where its documentation is weak.

Fine-tuning from feedback

After some months, logs contain many (question, chunk that helped, chunk that was retrieved but did not help) triples. Chunks that ranked high but did not help are hard negatives, the most useful training examples, because they teach the model the fine difference the base model misses. Fine-tuning a reranker on these is usually cheaper and safer than fine-tuning the embedding model, which requires re-embedding the whole corpus.

A real-life example

A bank's FAQ bot shows a spike in thumbs-down on questions containing "EMI" (equated monthly instalment) after a new loan product launches. The traces show the retriever returning the old personal-loan EMI page. The new product's documents use "monthly repayment" instead of "EMI", so keyword search misses them, and the embeddings rank the old page, which says "EMI" many times, higher.

Fixes: the product team adds "EMI" to the new documents' titles and FAQ; the ingestion adds a synonym field; 25 of the failed questions join the golden set. Two months later, the team has 8,000 labelled question-chunk pairs from feedback and fine-tunes their reranker on them. On the golden set, MRR rises, with the biggest gains on product-specific jargon.

Follow-up questions to expect

  • "Isn't thumbs-down data too sparse?" — Usually only a small share of users click. That is why you add implicit signals and weekly sampled reviews by people.
  • "Can you automatically update the system from feedback?" — Automatically add cases to a review queue, yes. Automatically change prompts or indexes without a human check, no; feedback can be wrong or malicious.
  • "How do you know a fix did not break something else?" — Run the full golden set before every release, not just the new case.