Scenario-Based AI Engineering Questions

Course Content

Scenario-Based AI Engineering Questions

26 sections · 146 lessons

Users ask multi-hop questions like: 'Which vendors mentioned in last quarter’s procurement docs also had delayed shipments?' How do you design retrieval systems for multi-document reasoning instead of single-chunk lookup?


The question is a join, so plan it like oneQ2 vendors thatalso shipped late?Hop 1: vendors inQ2 procurement PDFsHop 2: lateshipments inlogistics DBRetrieve PDFs,extract namesResolve Pvt Ltdvs Inc variantsSQL:delay_days above zero
The model extracts names from passages, which it does well; the intersection runs in code, which it does not.

What you need to know

Why one search cannot answer it

A single chunk never contains the answer. One document lists vendors; a different system records delays. Vector search finds chunks similar to the whole question, which are usually chunks that mention "vendors" and "delays" in general. Raising top-K to 50 only gives the model a pile of text and asks it to do a join in its head. Models are good at pulling entities out of a passage and unreliable at set operations over dozens of names.

Plan, extract, then compute

Python
plan = planner.decompose(question)     # returns typed hops, e.g.:# [{"id": "h1", "tool": "retrieve_extract", "query": "vendors in procurement docs",#   "filters": {"date": "2026-Q2"}, "schema": {"vendor": "str"}},#  {"id": "h2", "tool": "sql", "query": "vendors with delay_days > 0 in 2026-Q2"},#  {"id": "h3", "tool": "intersect", "inputs": ["h1", "h2"], "key": "vendor"}]from rapidfuzz import process, fuzzdef intersect(a: list[str], b: list[str], min_score=90):    matches = []    for name in a:        best = process.extractOne(normalise(name), [normalise(x) for x in b], scorer=fuzz.token_sort_ratio)        if best and best[1] >= min_score:            matches.append(name)    return matches

Hop 1 retrieves procurement documents and extracts vendor names into a list. Hop 2 queries the shipment database, which already holds delays as data. Hop 3 intersects in code. normalise lowercases and strips suffixes such as "Pvt Ltd" and "Inc."; fuzzy matching then handles small spelling differences.

  1. Decompose into typed hops with a planner prompt and a fixed output schema.
  2. Run each hop with its tool — retrieval plus extraction for documents, SQL for structured data.
  3. Resolve entities — "Acme Corp" and "ACME Inc." must become one vendor; show ambiguous matches instead of merging them silently.
  4. Compute joins, counts and dates in code.
  5. Answer with citations for every vendor, pointing at both sources.

Cap plans at about three hops. If a hop returns nothing, answer "insufficient evidence for hop 2" rather than building an answer from partial results.

The durable version

When the same entity types keep coming up, move the work to ingestion: extract entities and relations from every document into a graph or relational store. Retrieval then becomes vector search to find seed documents, graph traversal across documents, and a final chunk fetch for citations. This is the idea behind GraphRAG-style systems. Build it after you know which entities matter; building it first means extracting everything, at great cost.

Measure per-hop recall, exact match on a curated multi-hop question set, and cost, since multi-hop runs typically cost several single queries.

A real-life example

Scenario, numbers made up. A manufacturer's procurement team asks this exact question. The single-search assistant returns 5 vendors, of which 2 are wrong and 4 real ones are missing, because the relevant documents were spread across 60 PDFs.

With planning, hop 1 extracts 140 vendor names from Q2 procurement PDFs, and hop 2 returns 31 vendors with late shipments from the logistics database. Entity resolution merges "Shree Ganesh Castings Pvt Ltd" with "Shree Ganesh Castings" and flags two uncertain pairs for the user. The intersection gives 17 vendors, each cited to a PDF page and a shipment record. On a set of 80 multi-hop questions, exact-match accuracy rises from 22% to 64%.

Follow-up questions to expect

  • "When would you build a knowledge graph?" — When the same entity types and relationships are asked about often enough that extracting them once at ingestion is cheaper than per query.
  • "How do you stop the planner from inventing hops?" — A fixed schema of allowed tools, a hop cap, and a check that each hop's inputs exist.
  • "What does this cost?" — Several model calls per question instead of one. Route only questions the planner marks as multi-hop down this path.