Course Content
Advanced RAG
3 sections · 38 lessons
How does Vision-RAG handle image-based documents?
What you need to know
Why plain OCR fails
Many real documents carry meaning in their layout, not only their words:
- a table where the column header decides what a number means;
- a chart whose values exist only as bar heights;
- a form with ticked boxes, signatures and stamps;
- a two-column page that naive OCR reads straight across, mixing both columns into nonsense.
Once that structure is lost at ingestion, no chunking strategy or reranker can recover it.
Approach 1: the parsing pipeline
- Layout detection — find text blocks, tables, figures and headers on each page.
- Extract — OCR for text; table-structure recognition for tables (output as Markdown or HTML); a VLM writes a description of each chart or figure.
- Index — chunk and embed the extracted text as usual, with metadata such as page number and element type.
- Answer — send the text, and optionally the page image, to the model.
This keeps your normal text RAG stack. Modern document parsers increasingly use VLMs for the extraction step, which handles tables and messy scans far better than classic OCR.
Approach 2: native image retrieval (ColPali)
ColPali (2024) applies ColBERT's idea to images. A VLM turns each page image into many patch vectors; a text query becomes token vectors; MaxSim scores the match. There is no OCR at all. Retrieval returns whole page images, and a VLM reads them to answer. Newer models in the same family use stronger VLM backbones.
| Parsing pipeline | Native (ColPali-style) | |
|---|---|---|
| Ingestion | Complex; many tools | Simple; embed page images |
| Index size | Normal text index | Many vectors per page; large |
| Tables and charts | Only as good as the parser | Seen directly |
| Answering | Cheap text LLM possible | Needs a VLM on page images |
| Citations | Chunk or sentence | Page |
Costs to state honestly
Multi-vector image indexes are large. Sending page images to a VLM costs more tokens than sending a text chunk. Latency rises. Citations land on a page, not a sentence. For born-digital, well-structured documents — HTML help pages, Word files — a good parser and text RAG is cheaper and usually just as accurate.
A real-life example
A pharma company wants to search 15 years of regulatory submissions. Most older files are scanned PDFs. The questions are often about tables: "What was the assay result at 12 months, 30 °C / 65% RH, for batch B-114?"
With the text-only pipeline, OCR flattens the stability table. Values from different batches and time points end up on the same line, and the assistant sometimes quotes the 6-month value for the 12-month question — a serious error.
The team runs a trial on 2,000 scanned pages and 150 table questions written by the stability group:
- Parsing with a VLM-based table extractor: tables come out as Markdown grids with their headers. Most questions are answered correctly and cited to the page.
- ColPali-style retrieval plus a VLM: similar accuracy on tables, better on charts and handwritten annotations, but several times the cost per query.
They choose the parsing pipeline for the bulk of the archive and use native image retrieval only for a small collection of chart-heavy study reports. Every answer shows the page image next to the cited value so a scientist can check it.
Follow-up questions to expect
- "How do you handle a page with both text and a chart?" — In the pipeline approach, index the text chunk and the chart description as separate elements with the same page ID. In the native approach, the page is one unit.
- "How would you evaluate it?" — A question set focused on tables and figures, with the answer cell or value labelled. Measure retrieval of the right page and exact-value accuracy separately.
- "Can you combine both approaches?" — Yes: text retrieval and image retrieval as two lists fused with RRF, then a VLM answers from the top pages.