Scenario-Based AI Engineering Questions

Course Content

Scenario-Based AI Engineering Questions

26 sections · 146 lessons

Your AI assistant handles text well, but users now upload screenshots, diagrams, invoices, and videos. How do you design multimodal RAG systems that retrieve across text, images, and visual layouts?


What you need to know

Why one embedding space is not enough

Text embeddings cannot see a chart. Image embeddings such as CLIP are good at "a photo of a red car", but weak at reading the numbers on an invoice or the labels in an architecture diagram. Each content type needs its own processing, with one shared retrieval layer on top.

By content type

ContentIngestionSearched through
PDFs and scansLayout-aware parser keeps reading order, tables and figure boundariesText chunks, tables as markdown
Screenshots and diagramsVision-language model (VLM) writes a description; OCR for any textDescription and OCR text
Invoices and formsField extraction: vendor, date, line items, totalStructured fields plus text
Slides and visual pagesPage-image embeddings (ColPali style)Visual index
VideoSplit by scene; transcript chunks with timestamps plus keyframe descriptionsTranscript and descriptions; returns a timestamp

Layout-aware parsers include open-source tools such as Docling and Unstructured, and cloud document-AI services. ColPali-style models embed an image of the whole page as many small vectors and match them against query words, which works well for slides and visually dense pages.

Fusing the results

Text and visual indexes return scores on different scales, so combine them by rank, not score:

Python
def rrf(result_lists, k=60):    scores = {}    for results in result_lists:                 # each list is ordered best-first        for rank, doc_id in enumerate(results):            scores[doc_id] = scores.get(doc_id, 0) + 1 / (k + rank + 1)    return sorted(scores, key=scores.get, reverse=True)fused = rrf([bm25_hits, dense_hits, page_image_hits])

Reciprocal rank fusion (RRF) gives each document points for ranking high in any list. A reranker then orders the top candidates.

  1. Parse with layout — keep tables, figures and reading order.
  2. Create surrogates — descriptions, fields, markdown tables, each linked to page and region.
  3. Index — hybrid over surrogates; a page-image index for visual queries.
  4. Fuse and rerank — RRF, then a reranker.
  5. Generate with the original — give the VLM the retrieved crop as well as the text, and cite page plus region so users can check.

Costs and risks: VLM captioning is expensive, so do it once at index time, not per query. And a wrong caption quietly poisons the index, so check generated descriptions against OCR text from the same region.

A real-life example

Scenario, numbers made up. A finance-operations assistant answered only from text. Users start uploading vendor invoices, dashboard screenshots and architecture diagrams, and 40% of those questions fail.

The team adds a layout parser, invoice field extraction and VLM descriptions for images. Invoice questions ("What did we pay Kumar Logistics in March?") go from 35% to 89% correct, because totals are now fields, not scattered OCR text. Diagram questions improve less, from 30% to 62%, until a page-image index is added for "the slide with the payment flow". Captioning 200,000 images costs a one-time batch job; per-query cost barely changes.

Follow-up questions to expect

  • "Why not send every page image straight to a vision model at query time?" — Too slow and costly per query, and you still need retrieval to choose which pages to send.
  • "How do you evaluate it?" — Recall@k per content type, field-level accuracy on invoices, and whether the cited region really contains the answer.
  • "What about handwriting or poor scans?" — Route low-confidence OCR to a stronger model or a human, and mark answers from those pages as lower confidence.