AI Product Engineering: Shipping LLM Features That Last

Course Content

AI Product Engineering: Shipping LLM Features That Last

6 sections · 22 lessons

Retrieval when the model lacks your facts


In July, a Mumbai customer wrote: "bag was completely wet, rotis soggy, dal leaked". The v4 draft applied R2 and refunded 30%. The agent changed it to 50%. Mumbai has a monsoon rule, MUM-M2: from June to September, rain-damaged food gets 50% of the affected items' price, minimum ₹40. The model had never seen it.

TiffinGo's full refund policy is not seven rules. It is a 9-page base policy, 14 city annexes, and monthly promotion terms, about 60,000 tokens in total, and it changes about twice a month. The contract holds only the base rules. Four of the remaining v4 failures were city rules like MUM-M2.

This is the situation retrieval exists for: the model needs facts that are too many, or change too often, to live in the prompt. But retrieval is not the only answer, and for many teams it is not the best one.

From 60,000 policy tokens to the four sections that matterTicket, cityand order dateMetadatafilter: cityand valid datesBM25 over title,keywords, textTop 4 sections,about 900 tokensModel picksrule id;code prices itRecall at 4 rose from 8 of 11 to 11 of 11 after adding customer words like geela.
Most retrieval quality is decided before search runs — by splitting the policy into dated, city-tagged sections with words customers actually use.

Three ways to give the model facts

Put everything in the prompt. 60,000 tokens on every call is 72 million input tokens a day. At $2 per million input tokens and ₹88 to the dollar, that is about ₹12,700 a day, against about ₹750 a day for the whole feature today. Prompt caching cuts the price of repeated input sharply, perhaps to ₹1,300 a day, but there is a second problem. In a test on 20 Mumbai tickets with the full policy in the prompt, the model applied another city's rule 3 times. More text is more to confuse.

Filter by metadata. Every policy section belongs to a city (or to all cities) and has dates when it is valid. Filtering by the order's city and date leaves 6 to 40 sections, depending on the month. For most cities in most months, that is small enough to include all of it.

Search within the filtered sections. In festival months, some cities have 30 or more promotion sections. Then a search picks the 4 most relevant, about 900 tokens.

Everything in the prompt

  • Simplest; nothing to build
  • 60,000 tokens a call, about ₹1,300 a day even with caching
  • Model mixes up cities
  • Right choice when the policy is a few pages

Filter, then search

  • A small index and a filter to build
  • About 900 tokens a call
  • Only the right city and dates reach the model
  • Adds a step that can miss the right section

If your policy is five pages and changes twice a year, put it in the prompt and stop reading. Retrieval earns its complexity only when the facts are large, varied or changing.

Structure the source before you search it

Retrieval quality is mostly decided before any search runs. TiffinGo's policy was a set of long documents. The team split it into sections, one rule each, and gave each section metadata.

JSON
{  "id": "MUM-M2",  "title": "Mumbai monsoon: rain-damaged or soggy food",  "city": "Mumbai",  "valid_from": "2026-06-01",  "valid_to": "2026-09-30",  "keywords": "wet soggy rain leaked soaked geela bheeg baarish",  "text": "If food arrives wet, soggy or damaged by rain, refund 50% of the paid price of the affected units, minimum 40. Applies instead of R2 for these complaints.",  "pricing": {"percent": 50, "min_inr": 40}}

Three fields do most of the work. city, valid_from and valid_to let code filter exactly, so an expired Diwali promotion can never reach the model. keywords adds the words customers actually use, including Hindi ones like "geela" (wet), which never appear in the formal policy text. And pricing gives code the numbers to compute the amount, so the model still only chooses the rule.

Retrieval in code

Python
from datetime import datefrom rank_bm25 import BM25Okapidef tokens(text: str) -> list[str]:    return text.lower().replace(",", " ").replace(".", " ").split()def live_sections(sections: list[dict], city: str, on: date) -> list[dict]:    return [s for s in sections            if s["city"] in ("ALL", city)            and date.fromisoformat(s["valid_from"]) <= on <= date.fromisoformat(s["valid_to"])]def find_policy(sections: list[dict], ticket_text: str, city: str, on: date, k: int = 4) -> list[dict]:    live = live_sections(sections, city, on)    if len(live) <= k:        return live    docs = [tokens(f"{s['title']} {s['keywords']} {s['text']}") for s in live]    scores = BM25Okapi(docs).get_scores(tokens(ticket_text))    ranked = sorted(range(len(live)), key=lambda i: scores[i], reverse=True)    return [live[i] for i in ranked[:k]]

live_sections is the metadata filter, and it is exact. find_policy then ranks the survivors with BM25, a classic keyword-scoring method from the rank_bm25 package. Building a BM25 index over 40 sections on every request takes about a millisecond, so there is no index to keep in sync with policy edits.

Why keywords and not embeddings? For TiffinGo, customers' words and policy words overlap once the keywords field is added, and keyword search is easy to debug: you can see exactly which words matched. Embedding search is the better choice when phrasing varies more than a keyword list can cover, and it brings an embedding model and a vector store to run.

Put the sections into the contract

The retrieved sections go into the user message as a policy_sections list with their ids, titles and texts. Contract v5 adds three lines:

Text
POLICY SECTIONSpolicy_sections lists city or promotion rules that may apply to this order.If one applies, use its id as the line's rule; it replaces R2 when it says so.Use only rule ids from R1 to R3 or from policy_sections.

Code then extends the pricing table with the retrieved sections' pricing before calling price_draft, and the validator rejects any rule id that is not in R1 to R3 or in the retrieved list. The model can pick a city rule, but it cannot invent one.

Evaluate retrieval on its own

When a draft with retrieval is wrong, there are two possibilities: the right section was never retrieved, or it was retrieved and the model did not use it. The fixes are completely different, so measure them separately.

For each eval row whose correct answer uses a city section, check whether that section was in the retrieved list. This is recall at k: the share of rows where the needed section is in the top k. TiffinGo had 11 such rows. Before the keywords field, recall at 4 was 8 of 11; the misses were Hindi words and "leaked" versus "spilled". With keywords, it was 11 of 11. Only then did it make sense to judge the model's use of the sections.

Check your understanding

0 of 3 answered

1.Suppose TiffinGo's whole policy were only 4 pages and changed twice a year. What would be the best approach?

2.A Mumbai rain-damage ticket gets the wrong refund. The MUM-M2 section was not in the retrieved list. What should you fix?

3.Why does the validator reject rule ids that are not R1 to R3 or in the retrieved list?