Live Coding Interview Prep

Course Content

Live Coding Interview Prep

7 sections · 50 lessons

Build an end-to-end chatbot with memory, RAG, and evaluation pipeline.


Turn 2: 'and for COD?'Memory: turn1 aboutprepaid refundsRewrite:cash-on-deliveryrefundRetrieve:chunk c2 onlyAnswer cites[1], which is c2Guardrailpasses;save turnSearching the raw 'and for COD?' would have matched nothing.
The rewrite turns a follow-up into a searchable query, and the eval reports retrieval recall apart from groundedness.

What you need to know

This question checks whether you can integrate the earlier pieces cleanly. The parts:

partjobfrom which earlier question
Memorykeep recent turns and a summary per sessionsummary memory
Query rewrite"what about the blue one?" → a standalone querymulti-query and memory
Retrievertop-k chunks for the queryRAG, hybrid search
Generatoranswer only from numbered chunks, cite [n]citation tracking
Guardrailno valid citation → no answeroutput guardrails
Metricslatency per turnmetrics tracking
Evaluationrecall of the gold chunk, judged groundednessLLM-as-a-judge

Why rewrite the query? Retrieval sees only the text you give it. On turn 3, "and for COD orders?" retrieves nothing useful; "refund time for cash-on-delivery orders" does.

Why separate the metrics? A wrong answer has two possible causes. If the right chunk was not retrieved, fix chunking, embeddings or search. If it was retrieved but the answer is still wrong, fix the prompt or the model. One blended score cannot tell you which.

Python
import re, timefrom dataclasses import dataclassCITE = re.compile(r"\[(\d+)\]")SYSTEM = ("You are a support assistant. Answer only from the numbered context and cite "          "sources as [n]. If the context does not contain the answer, say so.")@dataclassclass Chunk:    id: str    text: str    source: strdef parse_citations(answer: str, chunks: list[Chunk]) -> list[dict]:    found, seen = [], set()    for m in CITE.finditer(answer):        n = int(m.group(1))        if 1 <= n <= len(chunks) and n not in seen:            seen.add(n)            found.append({"marker": n, "chunk_id": chunks[n - 1].id, "source": chunks[n - 1].source})    return foundclass Chatbot:    def __init__(self, retriever, llm_fn, memory_factory, metrics, top_k: int = 5) -> None:        self.retriever = retriever            # .search(query, k) -> list[Chunk]        self.llm_fn = llm_fn                  # (system, user) -> str        self.memory_factory = memory_factory  # () -> object with .turns, .add(), .context()        self.metrics, self.top_k = metrics, top_k        self.sessions: dict[str, object] = {}    def _standalone_query(self, memory, question: str) -> str:        if not memory.turns:            return question                   # first turn: nothing to resolve        try:            rewritten = self.llm_fn("Rewrite the final user question as a standalone search "                                    "query. Reply with the query only.",                                    f"{memory.context()}\nuser: {question}").strip()        except Exception:            rewritten = ""        return rewritten or question          # degrade to the raw question    def ask(self, session_id: str, question: str) -> dict:        started = time.perf_counter()        memory = self.sessions.setdefault(session_id, self.memory_factory())        query = self._standalone_query(memory, question)        chunks = self.retriever.search(query, self.top_k)        answer, citations = "I could not find anything about that.", []        if chunks:            context = "\n\n".join(f"[{n}] {c.text}" for n, c in enumerate(chunks, 1))            draft = self.llm_fn(SYSTEM, f"{memory.context()}\n\nContext:\n{context}\n\n"                                        f"Question: {question}")            citations = parse_citations(draft, chunks)            answer = draft if citations else "I could not answer that from our help articles."        memory.add("user", question)        memory.add("assistant", answer)        self.metrics.record("chat", (time.perf_counter() - started) * 1000)        return {"answer": answer, "citations": citations, "query": query,                "chunk_ids": [c.id for c in chunks]}

The evaluation harness (it reuses judge from the LLM-as-a-judge question, which returns None for an unparseable verdict):

Python
def evaluate(bot: Chatbot, cases: list, judge_fn) -> tuple[dict, list[dict]]:    """cases: objects with .question, .gold_chunk_id, .context, .reference."""    rows = []    for i, case in enumerate(cases):        out = bot.ask(f"eval-{i}", case.question)             # fresh session per case        verdict = judge(case, out["answer"], judge_fn)        rows.append({"question": case.question, "answer": out["answer"],                     "retrieved_gold": case.gold_chunk_id in out["chunk_ids"],                     "groundedness": verdict["groundedness"] if verdict else None})    graded = [r["groundedness"] for r in rows if r["groundedness"] is not None]    return {"n": len(rows),            "retrieval_recall": round(sum(r["retrieved_gold"] for r in rows) / len(rows), 3) if rows else None,            "groundedness": round(sum(graded) / len(graded), 2) if graded else None,            "unjudged": len(rows) - len(graded)}, rows

The tricky parts:

  • sessions.setdefault(...) creates memory for a new session on first use; the factory makes each session independent.
  • The rewrite has a fallback. If the rewrite call fails or returns nothing, the raw question is still a reasonable query; a failed helper call should not fail the turn.
  • The guardrail uses parsed citations. An answer whose markers point at no real chunk is treated like an answer with none.
  • A fresh session per eval case, so one case's memory cannot leak into the next and change its retrieval.

Complexity per turn: one rewrite call (after the first turn), one retrieval, one generation call; parsing is O(answer length). The evaluation costs one turn plus one judge call per case. Memory per session is bounded by the summary memory's budget.

A real-life example

A two-turn conversation with a keyword retriever, a scripted model, and the SummaryMemory and Metrics classes from earlier questions:

Python
DOCS = [Chunk("c1", "Refunds for prepaid orders take 5 working days.", "refunds.md"),        Chunk("c2", "Cash-on-delivery refunds are paid as store credit.", "refunds.md")]class KeywordRetriever:    def search(self, query, k):        words = {w for w in query.lower().split() if len(w) >= 4}    # skip "and", "for"        return [c for c in DOCS if words & set(c.text.lower().replace(".", "").split())][:k]script = iter(["Prepaid refunds take 5 working days [1].",              # turn 1 answer               "cash-on-delivery refund",                                # turn 2 rewrite               "They are paid as store credit [1]."])                   # turn 2 answerbot = Chatbot(KeywordRetriever(), lambda system, user: next(script),              memory_factory=lambda: SummaryMemory(lambda p: "summary", max_tokens=500),              metrics=Metrics())print(bot.ask("s1", "How long do prepaid refunds take?"))# {'answer': 'Prepaid refunds take 5 working days [1].', 'citations': [{'marker': 1,#  'chunk_id': 'c1', 'source': 'refunds.md'}], 'query': 'How long do prepaid refunds take?',#  'chunk_ids': ['c1', 'c2']}second = bot.ask("s1", "and for COD?")print(second["query"], second["chunk_ids"], second["answer"])# cash-on-delivery refund ['c2'] They are paid as store credit [1].print(bot.metrics.report("chat")["n"])                                  # 2
turnquery sent to retrievalretrievedmodel callsresult
1the raw question (no history yet)c1, c2 — both contain "refunds"1 (answer)cites [1] → c1
2"cash-on-delivery refund" (rewritten)c2 only2 (rewrite + answer)cites [1] → c2

Without the rewrite, turn 2 would have searched for "and for COD?"; the only word of four or more letters is "cod?", which is in neither chunk, so the bot would have said it found nothing.

This is the core loop of support chatbots at airlines, telecoms and banks; the evaluation harness is what lets their teams change a prompt on Tuesday and know by Wednesday whether retrieval recall or groundedness moved.

Follow-up questions to expect

  • "Two messages arrive for the same session at once?" — Put a per-session lock (in Redis for multiple servers) around ask, or the two turns interleave and memory records them out of order.
  • "How do you evaluate multi-turn behaviour?" — Add scripted conversations to the eval set: a sequence of turns with the expected standalone query and gold chunk at each step.
  • "What would you add next?" — Streaming, a re-ranker, feedback buttons wired into the eval set, and a sampled online judge on live traffic.