Course Content
Scenario-Based AI Engineering Questions
26 sections · 146 lessons
Your AI coding assistant works great on small repos but fails badly on monoliths with millions of lines of code. How do you build repository-scale context retrieval for coding agents?
What you need to know
Why small-repo tricks fail
In a small repo you can put most of the code in the prompt. A monolith with 4 million lines is tens of millions of tokens — far beyond any context window, and too slow and costly even if it fit. So retrieval must choose. Fixed-size text chunks cut functions in half, and embeddings are weak at rare identifiers like handleRetryPolicy, which is exactly what developers search for.
Four indexes over the same code
| Index | Built with | Answers |
|---|---|---|
| Structural | tree-sitter: one chunk per function, class or method, with path, signature and imports | "Show me CheckoutService.retry" |
| Graph | LSP or SCIP: definitions, references, calls, imports | "Who calls this? What does it implement?" |
| Lexical | ripgrep or BM25 | Exact identifiers, error strings, config keys |
| Semantic | Embeddings of symbol chunks with a one-line summary prepended | "Where do we handle rate limiting?" |
Structural chunking with the current py-tree-sitter API looks like this:
1import tree_sitter_python as tspython2from tree_sitter import Language, Parser34parser = Parser(Language(tspython.language()))56def symbol_chunks(path: str, source: bytes):7 tree = parser.parse(source)8 for node in tree.root_node.children:9 if node.type in ("function_definition", "class_definition"):10 name = node.child_by_field_name("name").text.decode()11 yield {"path": path, "symbol": name,12 "lines": (node.start_point[0] + 1, node.end_point[0] + 1),13 "text": source[node.start_byte:node.end_byte].decode()}Each chunk is a whole symbol with its file path and line numbers, so a function is never split. A real indexer also walks into classes for methods and handles decorated_definition nodes; tree-sitter has grammars for most languages, so the same idea works for Java or Go.
The retrieval flow
- Classify the query — a symbol name, an error message, or a question about behaviour.
- Exact first — symbol lookup and grep for identifiers and strings.
- Semantic next — embeddings for "where do we..." questions.
- Expand on the graph — one or two hops to callers, callees and interfaces.
- Rerank and pack — fit the best pieces into a fixed token budget, with paths and line numbers.
Two things matter even more than the indexes. A repo map — the directory tree with a one-line summary per module — gives the agent a mental model before it searches. And iterative, agentic search, where the agent calls grep, reads files and follows references step by step, beats one-shot retrieval clearly at monolith scale.
Keep the indexes fresh by re-indexing only the files changed in each commit: a full re-index takes hours, a delta takes seconds.
A real-life example
Scenario, numbers made up. An e-commerce company has a 4-million-line Java monolith. Its coding assistant uses 1,000-token text chunks and one embedding index. The team builds an eval from 200 merged PRs: given the linked issue, does retrieval return the files the PR touched in its top 10?
The baseline finds 31% of them. Symbol-level chunks raise that to 44%. Adding grep and the call graph, so "checkout retries twice" leads from the retry config to PaymentGatewayClient and its callers, reaches 61%. Giving the agent a repo map and letting it search in steps reaches 68%, and the share of the assistant's patches that compile on the first try rises with it.
Follow-up questions to expect
- "Why not use a model with a million-token context?" — The monolith is still far bigger, cost and latency grow with every token, and models find details less reliably in very long contexts. You still need to choose what goes in.
- "What tools would you give the agent?" —
grep,read_filewith line ranges,find_references,go_to_definition, and the repo map. They mirror how an engineer explores code. - "How do you keep embeddings current?" — Re-index changed files on each commit, keyed by file hash, and keep grep over the live checkout so the lexical path is never stale.