- MantraMindAI
- Blog
- Generative AI & LLMs
RAG in production: ingestion, context assembly, failure modes
Jai Rao
August 22, 202622 min read
The retriever is rarely what broke. A walk down the RAG pipeline - parsing, chunking, metadata, context budget, grounded prompts - and how to diagnose each failure.
A support assistant tells a customer that international returns are accepted within 60 days. The policy changed to 30 days seven months ago. Retrieval worked perfectly: the returns policy was the top hit, it was the right document, and the answer even carried a citation to it. What went wrong is that the ingestion job ran once, in March, and nobody wired up re-indexing. The model was faithful to its evidence, the evidence was stale, and every component in the pipeline reported success.
That is the shape of most real retrieval-augmented generation bugs. RAG gets discussed as though it were a retrieval technique, and the retriever is usually the part that is working. The failures live in the plumbing on either side: how documents became text, where the text got cut, what metadata travelled with each piece, how many pieces went into the prompt and in what order, what the model was told to do when the evidence ran out, and whether anyone ever measured the two halves of the system apart. This is a walk down that pipeline, stage by stage, with the failure that lives at each one. Retrieval itself gets treated the way you should treat it in code: a function you call that returns ranked chunks.
What changes when the answer has to come from a document
Ask a model a question with no evidence attached and the answer is assembled from whatever it absorbed during training. That answer has no timestamp, no source, and no scope. You cannot tell whether the model is reporting something it saw a thousand times or producing the most plausible continuation of your sentence, because both come out in the same confident register. You also cannot fix it. If the fact is wrong, or newly out of date, or something the asking user is not entitled to know, there is no lever short of retraining.
Attach retrieved passages and the failure profile changes shape. It does not shrink as much as the marketing suggests, but it becomes attributable. Every wrong answer now has a location. Either the passage that contains the answer was never retrieved, or it was retrieved and the model wrote something else anyway, or the passage itself was wrong. Those are three different bugs with three different owners, and you can tell them apart by looking at what went into the prompt. Debuggability, not accuracy, is the real payoff.
| Question you want to ask the system | Answering from weights | Answering from retrieved evidence |
|---|---|---|
| Where did this claim come from? | Unanswerable | A chunk id, a document, a section |
| How do I correct a wrong fact? | Retrain or patch with prompting | Fix the document, re-index |
| Can two users get different answers by entitlement? | No | Yes, if the ACL is on the chunk |
| Can I remove something permanently? | Not reliably | Delete it from the index |
| What happens when the knowledge is missing? | It guesses fluently | It can be made to decline |
Grounding also introduces failures a memory-only system cannot have. An index can go stale. An index can hand back a document the asking user was never meant to see. An index can contain two versions of the same policy that contradict each other. None of those are model problems, and none of them get better when you upgrade the model.
Ingestion is where these projects actually break
The stage that gets the least attention is the one that decides your ceiling. If the text going into the index is mangled, everything downstream is polishing.
Start with PDFs, because they are where the pain concentrates. A PDF is not a document format in the sense you want; it is a set of drawing instructions that place glyphs at coordinates. There are no paragraphs, no reading order, and no tables in the file — those are things a human eye reconstructs from layout. A naive text extractor hands you glyphs in the order the generator emitted them. For a two-column page that can mean the columns interleave line by line, producing text that is locally grammatical and globally nonsense. Running headers and footers land in the middle of sentences on every page boundary. Hyphenation splits words, ligatures come out as odd characters, and a scanned page contains no text at all, so you need OCR — and OCR on a fax-quality scan gets digits wrong, which matters a great deal when the document is a fee schedule.
Tables are the worst case, because a table's meaning is carried by its geometry. Flatten it and every cell loses its association with both its row and its column heading.
Table as printed: Region Return window Restocking fee Domestic 30 days none International 30 days 15%Flattened by a naive extractor: Region Return window Restocking fee Domestic 30 days none International 30 days 15%Emitted as row sentences: Returns, Domestic: return window 30 days, restocking fee none. Returns, International: return window 30 days, restocking fee 15%.The flattened form is worse than useless. Retrieved and handed to a model, it supports the claim that international returns carry no restocking fee just as readily as the correct claim. The row-sentence form costs you a few lines of parsing code and produces text that means exactly one thing. The rule generalises: never split a table across chunks, keep its caption and header row attached to every piece of it, and if the table is large enough to need splitting, repeat the header row.
Practical defences for the whole ingest stage. Use a layout-aware parser that reconstructs reading order and emits structure — headings, lists, tables — rather than one flat string. Detect and drop lines that repeat on nearly every page, which is what headers and footers are. Keep page numbers so citations can point somewhere. And fail loudly: a document that parsed to an empty string is worse than one that failed outright, because it sits in the index retrieving nothing and nobody notices for months. Assert a minimum character count per page and a sane ratio of alphabetic characters, and read five random chunks by eye every time you onboard a new source type. HTML exports arrive with cookie banners and navigation. Email threads carry the same quoted text five times. Spreadsheets have merged cells and sheets nobody mentioned. Every source type has its own way of being ugly, and you find out by looking.
Splitting text without cutting through the meaning
The default advice is to cut the extracted text every thousand characters. It is cheap, deterministic, and it will cut through the middle of sentences. That is not merely lossy — it manufactures false statements. Take a real policy sentence: Refunds are not available on items marked final sale, except where required by local law. Cut it after except where required and the first chunk now asserts, cleanly and quotably, a rule with its exception amputated. Retrieval will return it, the model will ground an answer in it, and the citation will point at a genuine document. Every trust signal in your product will fire on a wrong answer.
The second problem is reference. Prose leans on what came before it. A chunk containing only It must be requested within 30 days of delivery is retrievable by almost nothing and interpretable as almost anything, because it was defined two paragraphs up.
So split on structure first. Documents already carry boundaries that an author put there deliberately: headings, paragraphs, list items, table rows. Work down that hierarchy — take the largest unit that fits your budget, and only fall back to sentence boundaries, and then to raw characters, when a single paragraph is genuinely too big. Pack paragraphs into a chunk until the budget is spent, and never cross a heading, because a heading is the author telling you the subject just changed.
import redef chunk_by_structure(markdown, target=1200, overlap=300): """Pack paragraphs into chunks that never cross a heading.""" sections, path = [], [] for block in re.split(r"\n{2,}", markdown): block = block.strip() if not block: continue head = re.match(r"^(#{1,6})\s+(.+)$", block) if head: depth = len(head.group(1)) path = path[:depth - 1] + [head.group(2).strip()] sections.append((list(path), [])) elif sections: sections[-1][1].append(block) else: sections.append(([], [block])) chunks = [] for path, blocks in sections: buf = [] for block in blocks: if buf and sum(map(len, buf)) + len(block) > target: chunks.append({"section": " / ".join(path), "text": "\n\n".join(buf)}) buf = buf[-1:] if len(buf[-1]) <= overlap else [] buf.append(block) if buf: chunks.append({"section": " / ".join(path), "text": "\n\n".join(buf)}) return chunksTwo things to notice. The heading path is captured as data, not thrown away, so each chunk knows it came from Returns / International orders / Timeframe. And the overlap is one whole paragraph rather than a fixed slice of characters, so the carried text is always a complete thought. Character-window overlap is common and it reintroduces the exact problem you were trying to solve, just at the seam. A single paragraph longer than the target still comes out oversized, and that is the one case where you do drop to sentence boundaries.
Overlap is not free. It grows the index, and it makes retrieval return near-duplicates that consume slots in the prompt, which is why deduplication at assembly time is not optional once you use it. On size: small chunks retrieve more precisely, because the representation of a 150-word passage is about one thing while the representation of a whole 3,000-word section is about nothing in particular. Large chunks give the generator more to work with. The usual resolution is to have it both ways — index the paragraph, but send the surrounding section to the model. Retrieve small, expand before generation. And where your corpus has natural atoms — an FAQ entry, an API endpoint, a changelog line, a support ticket — chunk on those and ignore character counts entirely. A document that already comes in units does not need to be diced.
Every chunk travels with its paperwork
A chunk that is only text is nearly unusable in production. What travels alongside it decides whether you can filter, cite, refresh, or defend the system.
import hashlibdef to_record(chunk, doc): return { "chunk_id": hashlib.sha256( f"{doc['id']}|{chunk['section']}|{chunk['text']}".encode() ).hexdigest()[:16], # the heading path is part of what gets embedded, not just metadata "embed_text": f"{doc['title']} / {chunk['section']}\n\n{chunk['text']}", "text": chunk["text"], "doc_id": doc["id"], "title": doc["title"], "section": chunk["section"], "locator": f"{doc['url']}#page={chunk.get('page', 1)}", "effective_date": doc["effective_date"], "ingested_at": doc["ingested_at"], "acl": doc["acl"], # ["group:support", "group:legal"] "content_type": doc["type"], # policy | ticket | code | marketing }The embed_text line is the cheapest quality win in the whole pipeline. A chunk whose body says only 30 days from delivery, no restocking fee is not findable by a question about international refund timing. Prefix it with Returns policy / International orders / Timeframe and it is. You are giving the retriever back the context that chunking took away.
The dates do work too. A model has no idea what current means; it sees passages, not a calendar. If two versions of a policy sit in the index, the only way an answer can prefer the live one is for the metadata to say which is live — so the retriever can filter on it, and so the prompt can show effective dates and instruct the model to prefer the newest. Do both. ingested_at is separate and equally important: it is how you find out that one connector has been silently dead since March.
chunk_id is a hash of source, section, and text, which means it is stable across a re-index as long as the content has not changed. Citations that point at row numbers rot the first time you rebuild. And acl carries the entitlements of the source document down to the chunk, because that is the level at which the retriever will have to enforce them.
The context window is a budget, and you will overspend it
Retrieval hands you fifty ranked candidates. You can afford six. Everything about assembly follows from that.
Count tokens rather than guessing at them, and count all of them: the instructions, the question, the conversation history, the passages, and the space the answer itself needs. A configuration that fits at six passages will overflow the moment a user asks a follow-up and the history grows.
More evidence is not monotonically better. Past a handful of passages, the additions are mostly distractors, and the dangerous distractor is the one that is topically adjacent — a passage about domestic returns in an answer about international ones. The model has two plausible candidates and no principled basis for choosing. Teams routinely raise their passage count, watch quality fall, and blame the model. Measure it before you tune it.
Position matters as well. Attention across a long context is not uniform, and material buried in the middle of a large evidence block gets used less reliably than material at either end. The practical move is to keep your strongest passage adjacent to the question rather than in the middle of the pile, and to keep the instruction block in a fixed position so behaviour does not drift as the number of passages changes.
def assemble(hits, token_budget, count_tokens, per_doc=2): """Deduplicate, cap per document, fit the budget, then label.""" seen, per, kept = set(), {}, [] for hit in hits: key = " ".join(hit["text"].lower().split())[:400] if key in seen or per.get(hit["doc_id"], 0) >= per_doc: continue cost = count_tokens(hit["text"]) + 40 # 40 for the label line if cost > token_budget: continue # a smaller hit may still fit token_budget -= cost seen.add(key) per[hit["doc_id"]] = per.get(hit["doc_id"], 0) + 1 kept.append(hit) kept.sort(key=lambda h: h["score"], reverse=True) ordered = kept[1:][::-1] + kept[:1] # best passage last return [ f"[S{i+1}] {h['title']} / {h['section']} (effective {h['effective_date']})\n{h['text']}" for i, h in enumerate(ordered) ]The evidence block goes above the question in the prompt, so putting the top-scoring passage last places it next to what the user asked. Deduplication happens on normalised text, which catches both the near-duplicates your overlap created and the genuine article — the same paragraph living in a PDF, a wiki page, and an onboarding deck. Three copies of one paragraph read as corroboration to a model and cost you three slots. The per-document cap is crude and effective: it stops one verbose document from monopolising the context when the question spans several sources. Tighten it for narrow questions and expand a single hit to its neighbours instead; loosen it for broad ones like what changed in the returns policy this year.
Writing the answer without leaving the evidence
Three instructions do most of the work at generation time: answer only from the passages provided, cite the passage behind each factual claim, and if the passages do not contain the answer say so and stop. The third one is the instruction teams quietly delete, because it makes the demo look weaker. Keep it.
Abstention is a feature, not a degraded mode. An assistant that says the documents I can see do not cover international returns on gift orders; here are the two closest policies is more useful than one that guesses, because at the point of consumption a guess is indistinguishable from a fact. This has a product cost you should budget for: you need a real interface for no answer — show the near misses, offer the raw search results, offer a route to a human. Build that surface before you spend a week tuning the prompt, because without it your team will be under pressure to make the model answer anyway.
SYSTEM = """Answer using only the passages provided below.After each sentence containing a factual claim, cite the passage it camefrom as [S1], [S2] and so on. Cite only a passage that states the claim.If the passages do not contain the answer, reply with NOT_IN_SOURCESfollowed by one sentence naming what is missing. Never use outsideknowledge. If two passages disagree, say so and cite both."""CITE = re.compile(r"\[S(\d+)\]")def answer(question, retriever, user, client, count_tokens): hits = retriever.search(question, k=40, acl_filter=user.groups) passages = assemble(hits, 3000, count_tokens) if not passages: return {"status": "no_evidence", "answer": None, "cites": []} reply = client.complete( system=SYSTEM, user="\n\n".join(passages) + f"\n\nQuestion: {question}", ) if reply.startswith("NOT_IN_SOURCES"): return {"status": "declined", "answer": reply, "cites": []} used = {int(n) for n in CITE.findall(reply)} bad = {n for n in used if n < 1 or n > len(passages)} return {"status": "unverified" if bad else "answered", "answer": reply, "cites": sorted(used - bad)}Notice what that function returns: four distinct states, not a string. declined and no_evidence are different — one means the retriever found nothing the user may see, the other means it found passages that did not answer the question, and your interface should say different things. unverified means the model cited a passage number that was never sent, which happens more than you would like and is a signal, not an edge case.
That check is only structural, though. It proves the marker resolves to something you supplied; it says nothing about whether the passage supports the claim. For that you need a content check: compare the numbers, dates, and named entities in the sentence against the cited passage and flag a claim whose specifics appear nowhere in its source, or run a second, cheaper model call that judges each sentence against its citation. Either way, treat an unverified citation as unverified in the interface. A citation link is a trust signal and you should not hand one out for free.
One more thing that decides whether any of this matters: make the citation open the document at the right section, using the locator you stored at ingest. If verifying a claim costs the reader four minutes of scrolling through a 90-page PDF, nobody verifies anything and your citations are decoration.
Stale, leaked, contradictory, unsupported
Four failure modes account for most of the production incidents I have had to explain to somebody. They are worth naming individually because their fixes have nothing in common.
Stale. There are two mechanisms, and the second is nastier. Either nothing re-indexes, or updates are indexed but deletions are not. A withdrawn document that stays in the index retrieves confidently forever, because it is a real, well-written, professionally formatted document; nothing about it looks wrong. The fix is to treat ingestion as a synchronisation rather than an import: the pipeline needs a source of truth for what currently exists, a content hash per document so unchanged work is skipped, and explicit removal for anything that disappeared. Then alarm on the oldest ingested_at in the index, not the average. Averages look healthy while one connector has been broken for a month.
Leaked. Retrieval searches everything you indexed and knows nothing about who is asking unless you tell it. Two bugs are endemic. The first is filtering after retrieval: you fetch the top ten, drop the ones this user may not see, and hand over the remaining three — so the answer silently degrades, and the users with the least access get the worst results with no explanation. Worse, if you compensate by fetching more and refiltering, latency spikes for exactly those users. Permissions belong inside the retrieval query as a filter, resolved from the identity provider at request time. The second bug is capturing entitlements at ingest and never refreshing them, so the index still reflects a team someone left in April. And watch the derived-data trap: summaries, generated FAQ entries, and cached answers built from restricted documents inherit those restrictions, and almost nobody propagates the ACL into them. A leak here does not look like a leak. It looks like a helpful answer. Test it with a fixture user entitled to nothing, and assert that they get nothing.
Contradictory. Two passages say 30 days and 60 days. The model picks one — usually whichever reads more authoritatively — cites it correctly, and gives you no signal that a conflict existed. In order of value: keep one live version in the index, which is an organisational problem more than a technical one; failing that, filter on effective dates so only current documents are eligible; and instruct the model to report a disagreement rather than resolve it. That last output is genuinely valuable, by the way. These two documents disagree about the returns window is often a real problem in the documentation that nobody had noticed until a retrieval system put both paragraphs side by side.
Unsupported. The answer cites [S2], and [S2] is about something adjacent. This is the most corrosive failure because every artefact of trustworthiness is present. The usual causes: the model summarised across several passages and attached one label; the claim actually came from the model's own memory and the citation was ornamental; or the chunk was truncated so the qualifying clause never arrived. You cannot catch this by reading answers, because they read well. You catch it by sampling and scoring.
| What the user reports | Where to look first |
|---|---|
| Right answer, but for last year's policy | ingested_at on the cited chunk; whether deletions sync |
| Says it cannot find a page I can see in the wiki | The ACL filter, and whether that document parsed to non-empty text |
| Confidently wrong, cites a real document | Citation support checks, chunk truncation, two live versions |
| Good on general questions, bad on specific ones | Chunk size, and whether the heading path is in the embedded text |
| Quality dropped after we retrieved more | Distractors, and the strongest passage buried mid-context |
| Numbers in tables come out wrong | The parser: flattened tables, OCR on scans |
Two scores, because there are two failures
A single end-to-end quality number is the most expensive metric in RAG, because it cannot tell you which half is broken. Sixty percent end-to-end is equally consistent with a retriever that finds the right passage 95% of the time paired with a generator that ignores it, and with a scrupulously faithful generator that only received the right passage 60% of the time. Those two situations call for opposite work — one is a prompt and model problem, the other is a parsing, chunking, and indexing problem — and the aggregate score is silent on which one you have.
Split it. On the retrieval side, build a set of questions annotated with the chunk ids that actually contain the answer, then measure how often those arrive in the passages you sent. Recall is the number that dominates: if the evidence is not in the context, no amount of prompt work recovers it. Precision matters for a different reason — junk passages cost budget and add distractors — so track it, but do not trade recall away for it early on.
On the generation side, measure two separate things. Faithfulness asks whether every claim in the answer is supported by the passages that were supplied, which you assess by decomposing the answer into claims and checking each against the context. A model judge is the practical implementation; calibrate it against a couple of hundred human judgements before you trust its number, and re-calibrate when you change the judge. Correctness asks something different: whether the answer is right and complete against a reference. An answer can be perfectly faithful and badly incomplete, because it only received one of the three passages it needed.
The trick that makes these numbers actionable is conditioning. Score the generator only on the examples where retrieval succeeded. Otherwise you are marking the writer down for the librarian's mistakes, and your faithfulness metric moves every time somebody touches the chunker.
Two additions that catch things nobody else catches. Put questions in the eval set whose answers are genuinely absent from the corpus, and score whether the system correctly declined — without those, you cannot distinguish a helpful system from a compliant one, and every prompt change that raises helpfulness will quietly raise fabrication too. And keep the permission fixtures as assertions in the same suite, because access control is the one failure that is unacceptable rather than merely bad.
When retrieval is the wrong answer
RAG has become the reflexive architecture for anything involving documents, and a meaningful share of these projects should not exist. Four situations where something else fits better.
The corpus is small and fixed. If everything a question could need fits in the context window — one contract, one product manual, a policy handbook — put it all in the prompt. No chunk boundaries to get wrong, no index to keep fresh, no retrieval to debug. It costs more per call and it is slower, but with a stable prefix and prompt caching the arithmetic is friendlier than people expect, and the engineering you avoid is substantial. Reach for a pipeline when the corpus outgrows the window, not before.
The question is aggregate or relational. How many orders shipped late last quarter, by region is not a retrieval question. No passage contains the answer; the answer is a computation over rows. Point a parameterised query or a text-to-SQL step at the warehouse. Retrieval over a text dump of the same data will assemble a confident number out of whichever three rows came back, and it will be wrong in a way that looks exactly like being right. This is the most expensive mistake in the list.
You need behaviour, not facts. Fine-tuning is the right tool for how a model writes, what format it emits, what house style it follows, and what it refuses. It is a poor way to install facts: costly to update, impossible to cite, and the model cannot tell you where it learned anything. Facts belong in an index; manners belong in the weights.
The user wants documents. Sometimes the honest product is search with good snippets and no synthesis at all. For a domain expert who is going to read the source regardless, a generated paragraph is an extra layer to distrust. And if twenty questions cover most of your volume, write twenty answers, have a human approve them, and serve those — keep the pipeline for the tail.
The version of this system that works is unglamorous: a parser you have actually inspected, chunks that respect the boundaries the author wrote, metadata on every chunk, a small context assembled deliberately, a prompt that permits silence, and two scores you watch separately. Before you tune anything, check those six places. In my experience the expensive failure was already visible in one of them, weeks before anybody noticed the answers had gone wrong.