Course Content
Live Coding Interview Prep
7 sections · 50 lessons
Build an end-to-end chatbot with memory, RAG, and evaluation pipeline.
What you need to know
This question checks whether you can integrate the earlier pieces cleanly. The parts:
| part | job | from which earlier question |
|---|---|---|
| Memory | keep recent turns and a summary per session | summary memory |
| Query rewrite | "what about the blue one?" → a standalone query | multi-query and memory |
| Retriever | top-k chunks for the query | RAG, hybrid search |
| Generator | answer only from numbered chunks, cite [n] | citation tracking |
| Guardrail | no valid citation → no answer | output guardrails |
| Metrics | latency per turn | metrics tracking |
| Evaluation | recall of the gold chunk, judged groundedness | LLM-as-a-judge |
Why rewrite the query? Retrieval sees only the text you give it. On turn 3, "and for COD orders?" retrieves nothing useful; "refund time for cash-on-delivery orders" does.
Why separate the metrics? A wrong answer has two possible causes. If the right chunk was not retrieved, fix chunking, embeddings or search. If it was retrieved but the answer is still wrong, fix the prompt or the model. One blended score cannot tell you which.
1import re, time2from dataclasses import dataclass34CITE = re.compile(r"\[(\d+)\]")5SYSTEM = ("You are a support assistant. Answer only from the numbered context and cite "6 "sources as [n]. If the context does not contain the answer, say so.")78@dataclass9class Chunk:10 id: str11 text: str12 source: str1314def parse_citations(answer: str, chunks: list[Chunk]) -> list[dict]:15 found, seen = [], set()16 for m in CITE.finditer(answer):17 n = int(m.group(1))18 if 1 <= n <= len(chunks) and n not in seen:19 seen.add(n)20 found.append({"marker": n, "chunk_id": chunks[n - 1].id, "source": chunks[n - 1].source})21 return found2223class Chatbot:24 def __init__(self, retriever, llm_fn, memory_factory, metrics, top_k: int = 5) -> None:25 self.retriever = retriever # .search(query, k) -> list[Chunk]26 self.llm_fn = llm_fn # (system, user) -> str27 self.memory_factory = memory_factory # () -> object with .turns, .add(), .context()28 self.metrics, self.top_k = metrics, top_k29 self.sessions: dict[str, object] = {}3031 def _standalone_query(self, memory, question: str) -> str:32 if not memory.turns:33 return question # first turn: nothing to resolve34 try:35 rewritten = self.llm_fn("Rewrite the final user question as a standalone search "36 "query. Reply with the query only.",37 f"{memory.context()}\nuser: {question}").strip()38 except Exception:39 rewritten = ""40 return rewritten or question # degrade to the raw question4142 def ask(self, session_id: str, question: str) -> dict:43 started = time.perf_counter()44 memory = self.sessions.setdefault(session_id, self.memory_factory())45 query = self._standalone_query(memory, question)46 chunks = self.retriever.search(query, self.top_k)47 answer, citations = "I could not find anything about that.", []48 if chunks:49 context = "\n\n".join(f"[{n}] {c.text}" for n, c in enumerate(chunks, 1))50 draft = self.llm_fn(SYSTEM, f"{memory.context()}\n\nContext:\n{context}\n\n"51 f"Question: {question}")52 citations = parse_citations(draft, chunks)53 answer = draft if citations else "I could not answer that from our help articles."54 memory.add("user", question)55 memory.add("assistant", answer)56 self.metrics.record("chat", (time.perf_counter() - started) * 1000)57 return {"answer": answer, "citations": citations, "query": query,58 "chunk_ids": [c.id for c in chunks]}The evaluation harness (it reuses judge from the LLM-as-a-judge question, which returns None for an unparseable verdict):
1def evaluate(bot: Chatbot, cases: list, judge_fn) -> tuple[dict, list[dict]]:2 """cases: objects with .question, .gold_chunk_id, .context, .reference."""3 rows = []4 for i, case in enumerate(cases):5 out = bot.ask(f"eval-{i}", case.question) # fresh session per case6 verdict = judge(case, out["answer"], judge_fn)7 rows.append({"question": case.question, "answer": out["answer"],8 "retrieved_gold": case.gold_chunk_id in out["chunk_ids"],9 "groundedness": verdict["groundedness"] if verdict else None})10 graded = [r["groundedness"] for r in rows if r["groundedness"] is not None]11 return {"n": len(rows),12 "retrieval_recall": round(sum(r["retrieved_gold"] for r in rows) / len(rows), 3) if rows else None,13 "groundedness": round(sum(graded) / len(graded), 2) if graded else None,14 "unjudged": len(rows) - len(graded)}, rowsThe tricky parts:
sessions.setdefault(...)creates memory for a new session on first use; the factory makes each session independent.- The rewrite has a fallback. If the rewrite call fails or returns nothing, the raw question is still a reasonable query; a failed helper call should not fail the turn.
- The guardrail uses parsed citations. An answer whose markers point at no real chunk is treated like an answer with none.
- A fresh session per eval case, so one case's memory cannot leak into the next and change its retrieval.
Complexity per turn: one rewrite call (after the first turn), one retrieval, one generation call; parsing is O(answer length). The evaluation costs one turn plus one judge call per case. Memory per session is bounded by the summary memory's budget.
A real-life example
A two-turn conversation with a keyword retriever, a scripted model, and the SummaryMemory and Metrics classes from earlier questions:
1DOCS = [Chunk("c1", "Refunds for prepaid orders take 5 working days.", "refunds.md"),2 Chunk("c2", "Cash-on-delivery refunds are paid as store credit.", "refunds.md")]3class KeywordRetriever:4 def search(self, query, k):5 words = {w for w in query.lower().split() if len(w) >= 4} # skip "and", "for"6 return [c for c in DOCS if words & set(c.text.lower().replace(".", "").split())][:k]78script = iter(["Prepaid refunds take 5 working days [1].", # turn 1 answer9 "cash-on-delivery refund", # turn 2 rewrite10 "They are paid as store credit [1]."]) # turn 2 answer11bot = Chatbot(KeywordRetriever(), lambda system, user: next(script),12 memory_factory=lambda: SummaryMemory(lambda p: "summary", max_tokens=500),13 metrics=Metrics())14print(bot.ask("s1", "How long do prepaid refunds take?"))15# {'answer': 'Prepaid refunds take 5 working days [1].', 'citations': [{'marker': 1,16# 'chunk_id': 'c1', 'source': 'refunds.md'}], 'query': 'How long do prepaid refunds take?',17# 'chunk_ids': ['c1', 'c2']}18second = bot.ask("s1", "and for COD?")19print(second["query"], second["chunk_ids"], second["answer"])20# cash-on-delivery refund ['c2'] They are paid as store credit [1].21print(bot.metrics.report("chat")["n"]) # 2| turn | query sent to retrieval | retrieved | model calls | result |
|---|---|---|---|---|
| 1 | the raw question (no history yet) | c1, c2 — both contain "refunds" | 1 (answer) | cites [1] → c1 |
| 2 | "cash-on-delivery refund" (rewritten) | c2 only | 2 (rewrite + answer) | cites [1] → c2 |
Without the rewrite, turn 2 would have searched for "and for COD?"; the only word of four or more letters is "cod?", which is in neither chunk, so the bot would have said it found nothing.
This is the core loop of support chatbots at airlines, telecoms and banks; the evaluation harness is what lets their teams change a prompt on Tuesday and know by Wednesday whether retrieval recall or groundedness moved.
Follow-up questions to expect
- "Two messages arrive for the same session at once?" — Put a per-session lock (in Redis for multiple servers) around
ask, or the two turns interleave and memory records them out of order. - "How do you evaluate multi-turn behaviour?" — Add scripted conversations to the eval set: a sequence of turns with the expected standalone query and gold chunk at each step.
- "What would you add next?" — Streaming, a re-ranker, feedback buttons wired into the eval set, and a sampled online judge on live traffic.