Live Coding Interview Prep

Course Content

Live Coding Interview Prep

7 sections · 50 lessons

Build a logging and tracing system for LLM pipelines.


What you need to know

When a user reports "the bot gave a wrong answer", you need to see what happened inside that one request: what the retriever returned, which prompt version was used, what the model said and why it stopped. Tracing records that.

  • A trace is one request, identified by a trace_id.
  • A span is one stage inside it, with a span_id, a parent_span_id, a duration and attributes. Spans nest: rag_query contains retrieve and generate.
  • Structured logs are JSON objects rather than free text, so a log store can filter "all generate spans over 5 seconds with stop_reason = max_tokens".

contextvars holds values per execution context. Each asyncio task gets its own copy automatically, so two requests handled concurrently on one event loop never see each other's trace id — a module-level global would mix them. Threads are different: a new thread starts with an empty context. If you hand work to a thread pool, copy the context yourself with contextvars.copy_context().run(...), or use asyncio.to_thread, which does it for you.

Python
import contextvars, json, logging, time, uuidfrom contextlib import contextmanagerlog = logging.getLogger("llm")_trace_id: contextvars.ContextVar[str | None] = contextvars.ContextVar("trace_id", default=None)_stack: contextvars.ContextVar[tuple[str, ...]] = contextvars.ContextVar("span_stack", default=())@contextmanagerdef span(name: str, **attrs):    """Open a span; nested spans share the trace id and link to their parent."""    trace_token = _trace_id.set(uuid.uuid4().hex) if _trace_id.get() is None else None    parent = _stack.get()    span_id = uuid.uuid4().hex[:8]    stack_token = _stack.set(parent + (span_id,))    record = {"trace_id": _trace_id.get(), "span_id": span_id,              "parent_span_id": parent[-1] if parent else None, "name": name, **attrs}    started = time.perf_counter()    try:        yield record                              # the caller adds attributes as it learns them        record["status"] = "ok"    except Exception as exc:        record.update(status="error", error=f"{type(exc).__name__}: {exc}")        raise    finally:        record["duration_ms"] = round((time.perf_counter() - started) * 1000, 2)        log.info(json.dumps(record, default=str))        _stack.reset(stack_token)        if trace_token is not None:            _trace_id.reset(trace_token)

The tricky parts:

  • The outermost span creates the trace id; inner spans reuse it. trace_token remembers whether this span created it, so only that span clears it.
  • reset(token) restores exactly the previous value, which keeps nesting correct even when an exception unwinds several spans.
  • Logging in finally means a failed stage still produces a record, with status: error — the spans you most need are the failing ones.
  • time.perf_counter() is the right clock for durations: high resolution and never adjusted.

Complexity: O(1) work per span plus one log write; the log line grows with the number of attributes. Memory is O(depth of nesting) for the span stack.

A real-life example

A RAG request with a retrieval span and a failing generation span, captured with a list handler:

Python
records: list[dict] = []class ListHandler(logging.Handler):    def emit(self, r):        records.append(json.loads(r.getMessage()))log.addHandler(ListHandler())log.setLevel(logging.INFO)try:    with span("rag_query", user_id="u_123"):        with span("retrieve", top_k=5) as s:            s["n_chunks"] = 5        with span("generate", model="claude-opus-5", prompt_version=7) as s:            s["input_tokens"] = 1840            raise TimeoutError("provider timeout")except TimeoutError:    passfor r in records:    print(r["name"], r["status"], r["parent_span_id"] is None, r["trace_id"] == records[0]["trace_id"])# retrieve ok False True# generate error False True# rag_query error True True
order loggedspanparentstatuswhy this order
1retrieverag_queryokinner spans finish first
2generaterag_queryerrorexception recorded, then re-raised
3rag_querynone (root)errorthe exception passed through it

All three share one trace_id, so a single query in the log store shows the whole request, including the 1,840 input tokens sent before the timeout.

This is what lets an on-call engineer at a food-delivery company answer "why did the bot tell this customer their refund was rejected?" in minutes: find the trace, see the retrieved chunks, see the prompt version.

Follow-up questions to expect

  • "What would you log for each generation?" — Model, prompt name and version, input and output tokens, cost, latency and time to first token, stop_reason, cache hit, and the ids of retrieved chunks.
  • "Should you log full prompts and answers?" — They contain personal data. Redact, sample, restrict access, and keep them for a shorter time than the metrics.
  • "What would you use in production?" — OpenTelemetry with its generative-AI semantic conventions, exported to your tracing backend, or an LLM-specific tool such as Langfuse or LangSmith. The code above is the shape they record.