Advanced RAG

Course Content

Advanced RAG

3 sections · 38 lessons

How do you handle multiple programming languages in a code-focused RAG system?


What you need to know

Why text-RAG defaults fail on code

  • Character-based chunks cut functions in half and separate a function from its doc comment.
  • Identical names recur everywhere: main, handler, config, utils. Without the file path, a chunk is ambiguous.
  • Queries mix natural language ("where do we retry failed payments?") with exact tokens (ECONNRESET, get_user_by_id, PAYMENTS_MAX_RETRY).
  • Understanding one function often requires its callers, the types it uses, and its tests.

Chunk by syntax

LangChain's language-aware splitter uses separators such as func and class for each language. It is better than plain text splitting, but it is still a text heuristic:

Python
from langchain_text_splitters import Language, RecursiveCharacterTextSplittergo_src = '''package ledger// PostEntry writes a double-entry record.func PostEntry(tx Tx, e Entry) error {    if e.Amount <= 0 {        return ErrInvalidAmount    }    return tx.Insert(e)}// Reverse creates the opposite entry.func Reverse(tx Tx, id string) error {    e, err := tx.Get(id)    if err != nil {        return err    }    e.Amount = -e.Amount    return tx.Insert(e)}'''splitter = RecursiveCharacterTextSplitter.from_language(    Language.GO, chunk_size=200, chunk_overlap=0)for doc in splitter.create_documents([go_src]):    print("---\n" + doc.page_content)

Run it and look closely: each function gets its own chunk, but the // Reverse creates… comment ends up at the bottom of the PostEntry chunk, and PostEntry's own comment sits with the package line. The splitter cut before func, not before the comment that belongs to it. A real parser such as tree-sitter builds the syntax tree, so you can take each function node together with its leading comment. For a production system across many languages, parse; don't split.

Metadata and context per chunk

  • language, repo, path, symbol, kind (function, class, test, config, migration), commit SHA.
  • A contextual header: file path, enclosing class or module, and relevant imports.
  • Filtering on language or repo when the question names one removes most cross-language noise — the cheapest win available.

Retrieval that suits code

  • Hybrid search. BM25 nails exact identifiers and error strings; dense search handles "where do we…" questions. Fuse with RRF.
  • Multi-representation. Index the code, its doc comment, and an LLM-written English description ("Posts a ledger entry; rejects non-positive amounts"), all pointing at the same chunk. English questions then match code written in any language.
  • Code-aware embeddings. Models trained on code and natural-language pairs do much better than general text models.
  • Graph expansion. Store import, call and "tested by" edges. After retrieving a function, pull in its definition dependencies or callers when the question needs them. Multi-hop retrieval pays off in code more than almost anywhere.

Watch for boilerplate

Generated clients, vendored libraries and scaffold files are near-identical across services and crowd out real results. Exclude vendored and generated paths at ingestion, and deduplicate by content hash.

A real-life example

A fintech's engineering-wiki assistant is extended to answer questions about code in a monorepo with Go services, Python data jobs, TypeScript frontends and Terraform. An engineer asks: "Where do we reject negative ledger amounts, and which service calls it?"

The first version used plain 1,000-character chunks and a general embedding model. It returned a TypeScript form validator (it contained "amount" and "negative") and half of the Go function, without its name.

The rebuilt pipeline:

  • tree-sitter chunks per function for all four languages, each with path and package header;
  • BM25 plus a code embedding model, fused with RRF;
  • an English description per function, generated once at ingestion and refreshed when the function's hash changes;
  • a call graph from the build system.

Now PostEntry ranks first via its description ("rejects non-positive amounts"), and graph expansion adds its two callers in the payments-api service. The answer names the file, the function, and the callers, each with a link to the exact line at the current commit. The team evaluates on 150 questions written by engineers, split by language.

Follow-up questions to expect

  • "How do you keep it fresh with constant commits?" — Index on merge to the main branch, re-parse only changed files, and upsert by symbol ID and content hash.
  • "How do you handle very large functions?" — Split at inner blocks but prefix each piece with the function signature, so every piece still says what it belongs to.
  • "Should the assistant search code or call tools?" — Both. A language server or code-search tool can answer "find all references" exactly; RAG handles "how does this work" questions.