Applied AI Engineering: From Prompt to Production

Course Content

Applied AI Engineering: From Prompt to Production

9 sections · 29 lessons

The RAG pipeline end to end


PolicyPal now has a solid prompt and a JSON contract. It still knows nothing about Harbourline. The facts live in 400 PDFs: 230 HR policies and 170 IT standards, about 4,800 pages and 2.6 million tokens. As you calculated in Section 1, that is far too much to send with every question, both in cost and in context size.

Retrieval-augmented generation, or RAG, solves this with a simple idea. For each question, find the few passages most likely to contain the answer, and put only those in the prompt. The model's job shrinks from "know Harbourline's policies" to "answer from these five passages", which, as the "12 days" story showed, is a job it does well.

In this lesson you build the simplest RAG pipeline that works end to end, and you measure it the right way: retrieval first, on its own, because if the right passage never reaches the prompt, no prompt can save the answer.

Answering one question from 400 PDFsQuestion from aPune employeeEmbed withbge-small,384 numbersKeep onlyIndia andGlobal chunksTop 5 bycosinesimilarityAnswer fromthose 5,with citationsNaive recall@5 was 0.62; answers were right 92% of the time when the chunk arrived.
If the right chunk never reaches the prompt, no prompt can fix the answer — so measure retrieval on its own.

Two pipelines, not one

RAG is really two pipelines that share an index. The indexing pipeline runs offline, whenever policies change. The answering pipeline runs online, on every question.

Indexing (offline, nightly)

  • Load each PDF and extract text per page
  • Split text into chunks of a few hundred tokens
  • Turn each chunk into a vector with an embedding model
  • Store vectors, text and metadata together

Answering (online, per question)

  • Turn the question into a vector with the same model
  • Find the chunks whose vectors are closest
  • Put the top chunks into the prompt template
  • Generate the answer, with citations

An embedding model turns a piece of text into a list of numbers, a vector, such that texts with similar meaning get vectors that point in similar directions. PolicyPal uses BAAI/bge-small-en-v1.5, a small open model that runs on a CPU, produces 384 numbers per text, and reads at most 512 tokens of input. Keep that last number in mind; it matters in the next lesson.

Indexing: load, split, embed, store

The first version splits text into fixed windows of 500 tokens, counted with the embedding model's own tokenizer. It is deliberately naive, so you have a baseline to improve.

Python
# policypal/ingest.py  (version 1: naive fixed-size chunks)import jsonfrom pathlib import Pathimport numpy as npfrom pypdf import PdfReaderfrom sentence_transformers import SentenceTransformerembedder = SentenceTransformer("BAAI/bge-small-en-v1.5")tok = embedder.tokenizerdef pdf_pages(path: Path) -> list[tuple[int, str]]:    reader = PdfReader(path)    return [(i + 1, page.extract_text() or "") for i, page in enumerate(reader.pages)]def fixed_chunks(text: str, size: int = 500) -> list[str]:    ids = tok(text, add_special_tokens=False)["input_ids"]    return [tok.decode(ids[i:i + size]) for i in range(0, len(ids), size)]def build_index(pdf_dir: str, meta: dict, out: str = "index") -> None:    chunks = []    for path in sorted(Path(pdf_dir).glob("*.pdf")):        info = meta[path.name]          # title, country, version, effective_from        for page, text in pdf_pages(path):            for piece in fixed_chunks(text):                chunks.append({**info, "file": path.name, "page": page, "text": piece})    vecs = embedder.encode([c["text"] for c in chunks], batch_size=64,                           normalize_embeddings=True, show_progress_bar=True)    Path(out).mkdir(exist_ok=True)    np.save(f"{out}/vectors.npy", vecs.astype(np.float32))    Path(f"{out}/chunks.json").write_text(json.dumps(chunks))

Two things here are more important than they look. The meta dictionary comes from the policy register HR already maintains: title, country (India, UK or Global), version and effective date. That metadata is what later lets you filter by country and ignore superseded versions. And vectors are normalised to length 1, so a plain dot product equals cosine similarity.

This naive index held about 8,900 chunks, many of them short page tails of a few dozen tokens. Embedding them took about four minutes on a laptop CPU. The vectors take 8,900 × 384 × 4 bytes, about 14 MB, which fits comfortably in memory.

Answering: search, then generate

At question time, embed the question with the same model and take the closest chunks. The BGE models recommend a short instruction in front of search queries, and it measurably helps, so PolicyPal adds it.

Python
# policypal/retrieve.py  (version 1: dense search only)import jsonfrom pathlib import Pathimport numpy as npfrom policypal.ingest import embedderQUERY_PREFIX = "Represent this sentence for searching relevant passages: "class DenseIndex:    def __init__(self, path: str = "index"):        self.vecs = np.load(f"{path}/vectors.npy")        self.chunks = json.loads(Path(f"{path}/chunks.json").read_text())        self.country = np.array([c["country"] for c in self.chunks])    def search(self, question: str, country: str, k: int = 5) -> list[dict]:        q = embedder.encode([QUERY_PREFIX + question], normalize_embeddings=True)[0]        scores = self.vecs @ q        allowed = (self.country == country) | (self.country == "Global")        scores = np.where(allowed, scores, -np.inf)       # filter before ranking        top = np.argsort(-scores)[:k]        return [self.chunks[i] | {"score": float(scores[i])} for i in top]

The country filter is applied before ranking, by setting disallowed scores to minus infinity. That single line fixes the wrong-country answers from Section 1 at the source: a Pune employee's search can no longer return a UK-only chunk, however similar it looks. Filtering after ranking would be weaker, because the top five might all be UK chunks and nothing would be left.

The answering function then connects the pieces you already have: DenseIndex.search, the build_messages template from Section 2, and the ask function that validates the JSON reply.

A brute-force dot product over 8,900 vectors takes about a millisecond. You do not need a vector database for this corpus. You would want one with millions of vectors, frequent updates, several replicas, or complex filters and access control; if you already run Postgres, the pgvector extension is often the least new infrastructure. Choose the storage when the numbers require it, not before.

Measure retrieval on its own

A RAG answer can be wrong for two very different reasons: the right passage was never retrieved, or it was retrieved and the model misused it. You fix these in different places, so measure them separately.

For retrieval, HR helped label a retrieval check set of 60 real questions, each with the file and section that answers it. The metric is recall@5: the fraction of questions for which at least one correct chunk appears in the top five results.

Python
# policypal/evals/retrieval.pydef hit(chunk: dict, gold: dict) -> bool:    return chunk["file"] == gold["file"] and gold["section"] in chunk["text"]def recall_at_k(index, cases: list[dict], k: int = 5) -> float:    found = 0    for case in cases:        results = index.search(case["question"], case["country"], k=k)        found += any(hit(c, case["gold"]) for c in results)    return found / len(cases)

The naive pipeline scored 0.62: for 23 of the 60 questions, the answer was not in the top five. The team read every miss and sorted them by cause. Nine had the answer split across a chunk boundary. Six were questions about tables whose text came out of the PDF scrambled. Five used an exact term, such as "Form 12BB" or "ESPP", that semantic search did not match. Three found an old version of a policy. Each group points to a specific fix, and the next three lessons take them in order.

Check your understanding

0 of 3 answered

1.PolicyPal applies the country filter before ranking rather than after. Why does the order matter?

2.The naive pipeline's answers were correct 92% of the time when the right chunk was retrieved, and 71% overall. Where should the team work next?

3.When would PolicyPal need a dedicated vector database instead of a NumPy array?