Advanced RAG

Course Content

Advanced RAG

3 sections · 38 lessons

When should knowledge graphs be used instead of RAG?


Similarity versus traversalVector RAG answers well• Which passage explains re-KYC rules?• Paraphrased, fuzzy questions• Single facts inside documents• Cheap to build and keep freshA graph answers well• Clients two hopsfrom a sanctioned entity• Counts over every matching contract• Themes across 20,000 complaints• Completeness: all matches, not top 5
Top-k can never promise it found every connected company; a graph query can, which is why compliance questions go there.

What you need to know

What vector search cannot do

Vector search answers "which passages look like this question?" It cannot:

  • Follow edges several times — company → director → other company → sanctions list.
  • Aggregate — "how many of our contracts include this indemnity clause?" needs every match counted, not the top 5.
  • Answer global questions — "what are the recurring themes across all complaints?" has no single chunk that contains the answer.
  • Guarantee completeness — top-k returns some matches; a graph query returns all of them.

Domains where relationships are the data suit graphs: ownership and director networks, drug interactions, service dependency maps, org charts, entitlements.

GraphRAG: the practical hybrid

Microsoft's GraphRAG (2024) popularised combining the two:

  1. Extract — an LLM reads every chunk and writes out entities, relationships and short descriptions.
  2. Cluster — a community-detection algorithm (Leiden) groups densely connected entities into communities, at several levels.
  3. Summarise — an LLM writes a summary of each community.
  4. Local search — for entity questions, start at the entities named in the query and walk their neighbourhood, pulling linked text chunks.
  5. Global search — for "big picture" questions, map the question over community summaries and reduce the partial answers into one.

Later variants reduce the cost: LazyGraphRAG skips the up-front LLM summarisation and does more work at query time; LightRAG uses a lighter graph index alongside vectors.

The costs

  • Indexing — an LLM call over every chunk, which is large for big corpora.
  • Schema and quality — extracted entities need deduplication ("Sagar Traders", "Sagar Traders Pvt Ltd", "M/s Sagar Trading") and the schema needs an owner.
  • Freshness — a changed document can change entities, edges and community summaries. Much harder than upserting a vector.

When the relationships already exist in structured form — a company registry, a service catalogue — load them directly into a graph. LLM extraction is for relationships hidden in text.

A real-life example

A bank's compliance team needs to know: "Which of our corporate borrowers share a director or a major shareholder with any entity on the latest sanctions list?"

A vector search over KYC documents returns passages that mention sanctions or directors, but it cannot connect Company A's director to Company B's shareholding to a sanctions entry. Missing one link is a regulatory failure, and top-k gives no guarantee of completeness.

The team builds a graph from structured sources — the corporate registry data they already license, their own KYC records, and the sanctions list — with nodes for companies and people, and edges for DIRECTOR_OF and SHAREHOLDER_OF. The question becomes a two-hop graph query that returns every match with the path that connects it. The assistant then uses ordinary RAG over KYC documents to explain each match in plain language, citing the documents.

Separately, for "What are the most common reasons our customers raise complaints about loan closure?", they run GraphRAG-style global search over 20,000 complaint notes, because the answer is a pattern across the corpus, not a fact in one note.

Follow-up questions to expect

  • "When is GraphRAG overkill?" — When questions are mostly single-fact lookups from documents. Hybrid search plus a reranker is cheaper and easier to maintain.
  • "How would you keep the graph fresh?" — Re-extract only changed documents, merge entities by stable IDs where they exist, and rebuild community summaries on a schedule rather than on every change.
  • "How do you let an LLM query a graph safely?" — Text-to-Cypher with a read-only role and a small, documented schema; or better, a set of pre-written parameterised queries exposed as tools.