Course Content
Scenario-Based AI Engineering Questions
26 sections · 146 lessons
Your retrieval pipeline returns technically correct documents — but they’re written in legal language users can’t understand. How do you make RAG systems retrieve answers optimized for comprehension, not just relevance?
What you need to know
Why the retriever picks hard text, and why that is not the whole problem
Users write "Is my knee surgery covered in the first year?" The policy says "Expenses attributable to pre-existing diseases shall be excluded until the expiry of 36 months of continuous coverage." The words hardly overlap, so plain questions match plain-sounding chunks, not always the right clause. And when the right clause is found, the user still cannot read it.
Trying to fix both with one retriever fails. Rank for simplicity and you lose the correct clause; rank for correctness and you keep the unreadable text. So split the jobs.
Index time
1for clause in parse_policy(pdf):2 gloss = small_llm(f"Explain in plain English what this clause means for a customer, "3 f"in 2 sentences. Keep every number and date exactly:\n{clause.text}")4 for kind, text in (("verbatim", clause.text), ("gloss", gloss)):5 index.upsert(id=f"{clause.id}:{kind}", vector=embed(text),6 payload={"clause_id": clause.id, "kind": kind, "text": text,7 "doc": clause.doc, "section": clause.section})Both vectors carry the same clause_id. At query time, collapse hits by clause_id, so a clause found through its gloss is still grounded in its original words.
Answer time
- Plain answer first — at a stated reading level, such as "a 14-year-old can follow it".
- Quote the clause — the exact text, with document and section, under the answer.
- Protect key terms — never paraphrase amounts, dates, time periods, liability caps, "indemnify", or "shall" versus "may". That is where simplifying changes the legal meaning.
- Check entailment — an NLI model or a small LLM judge confirms the plain answer follows from the quoted clause; if not, block it and show only the clause.
What to measure
| Metric | Why |
|---|---|
| Retrieval recall | Must not drop when you add glosses |
| Readability grade | Checks the plain answer is actually plain |
| Faithfulness rate | Simplification must not change meaning |
| Follow-up rate | The real signal: did the user need to ask again? |
A real-life example
Scenario, numbers made up. A health insurer's policy assistant answers from 40 policy documents. 38% of conversations have a follow-up like "what does that mean?", and support tickets quote the bot's answers back with confusion.
The team generates glosses for 6,000 clauses and switches to the plain-first template. Recall@5 on their 300-question set rises from 72% to 81%, because plain questions now match glosses. In the A/B test, the follow-up rate falls from 38% to 17%. The entailment check blocks about 2% of answers — most often where the gloss had turned "up to 36 months" into "after 3 years", a small change with a real legal difference.
Follow-up questions to expect
- "Why not simplify only at generation time?" — Then retrieval still misses clauses because users' words do not match legal wording. The gloss helps retrieval as much as reading.
- "What if the gloss itself is wrong?" — Check glosses once at index time with the same entailment test, and have a domain expert review the highest-traffic clauses.
- "What about users who read Hindi or Tamil?" — Generate the plain answer in the user's language, but keep the quoted clause in the legally binding language.