Course Content
Scenario-Based AI Engineering Questions
26 sections · 146 lessons
Your AI assistant handles text well, but users now upload screenshots, diagrams, invoices, and videos. How do you design multimodal RAG systems that retrieve across text, images, and visual layouts?
What you need to know
Why one embedding space is not enough
Text embeddings cannot see a chart. Image embeddings such as CLIP are good at "a photo of a red car", but weak at reading the numbers on an invoice or the labels in an architecture diagram. Each content type needs its own processing, with one shared retrieval layer on top.
By content type
| Content | Ingestion | Searched through |
|---|---|---|
| PDFs and scans | Layout-aware parser keeps reading order, tables and figure boundaries | Text chunks, tables as markdown |
| Screenshots and diagrams | Vision-language model (VLM) writes a description; OCR for any text | Description and OCR text |
| Invoices and forms | Field extraction: vendor, date, line items, total | Structured fields plus text |
| Slides and visual pages | Page-image embeddings (ColPali style) | Visual index |
| Video | Split by scene; transcript chunks with timestamps plus keyframe descriptions | Transcript and descriptions; returns a timestamp |
Layout-aware parsers include open-source tools such as Docling and Unstructured, and cloud document-AI services. ColPali-style models embed an image of the whole page as many small vectors and match them against query words, which works well for slides and visually dense pages.
Fusing the results
Text and visual indexes return scores on different scales, so combine them by rank, not score:
1def rrf(result_lists, k=60):2 scores = {}3 for results in result_lists: # each list is ordered best-first4 for rank, doc_id in enumerate(results):5 scores[doc_id] = scores.get(doc_id, 0) + 1 / (k + rank + 1)6 return sorted(scores, key=scores.get, reverse=True)78fused = rrf([bm25_hits, dense_hits, page_image_hits])Reciprocal rank fusion (RRF) gives each document points for ranking high in any list. A reranker then orders the top candidates.
- Parse with layout — keep tables, figures and reading order.
- Create surrogates — descriptions, fields, markdown tables, each linked to page and region.
- Index — hybrid over surrogates; a page-image index for visual queries.
- Fuse and rerank — RRF, then a reranker.
- Generate with the original — give the VLM the retrieved crop as well as the text, and cite page plus region so users can check.
Costs and risks: VLM captioning is expensive, so do it once at index time, not per query. And a wrong caption quietly poisons the index, so check generated descriptions against OCR text from the same region.
A real-life example
Scenario, numbers made up. A finance-operations assistant answered only from text. Users start uploading vendor invoices, dashboard screenshots and architecture diagrams, and 40% of those questions fail.
The team adds a layout parser, invoice field extraction and VLM descriptions for images. Invoice questions ("What did we pay Kumar Logistics in March?") go from 35% to 89% correct, because totals are now fields, not scattered OCR text. Diagram questions improve less, from 30% to 62%, until a page-image index is added for "the slide with the payment flow". Captioning 200,000 images costs a one-time batch job; per-query cost barely changes.
Follow-up questions to expect
- "Why not send every page image straight to a vision model at query time?" — Too slow and costly per query, and you still need retrieval to choose which pages to send.
- "How do you evaluate it?" — Recall@k per content type, field-level accuracy on invoices, and whether the cited region really contains the answer.
- "What about handwriting or poor scans?" — Route low-confidence OCR to a stronger model or a human, and mark answers from those pages as lower confidence.