LLMOps & Deployment

Course Content

LLMOps & Deployment

6 sections · 40 lessons

What logging and tracing strategies help debug LLM pipelines?


The trace behind "it suggested a function that does not exist"request 4.3 sretrieve 0.21 sLLM call 3.9 s8 chunks, old branchrerank kept 36,120 in, 410 outfinish: stop
The model used what it was given; the stale chunks in the retrieval span point at the indexing job, not the prompt.

What you need to know

What to put on the LLM span

  • Model and snapshot ID; prompt version; temperature and max_tokens
  • Input and output tokens; computed cost
  • TTFT and total latency; finish_reason (stop, length, tool_calls)
  • The prompt and response text (redacted), or a pointer to where they are stored
  • Your own attributes: tenant, route, flag arm, index version

OpenTelemetry GenAI conventions

OTel defines standard attribute names for model calls. They are still marked "Development" in 2026, but the core names are stable in practice and supported by Langfuse, Arize Phoenix, LangSmith and most APM vendors.

Python
from opentelemetry import tracefrom opentelemetry.sdk.trace import TracerProviderfrom opentelemetry.sdk.trace.export import ConsoleSpanExporter, SimpleSpanProcessorprovider = TracerProvider()provider.add_span_processor(SimpleSpanProcessor(ConsoleSpanExporter()))trace.set_tracer_provider(provider)tracer = trace.get_tracer("code-assistant")with tracer.start_as_current_span("chat model-a-2026-03") as span:    span.set_attribute("gen_ai.operation.name", "chat")    span.set_attribute("gen_ai.provider.name", "openai")    span.set_attribute("gen_ai.request.model", "model-a-2026-03")    span.set_attribute("gen_ai.request.max_tokens", 800)    # ... call the model here ...    span.set_attribute("gen_ai.usage.input_tokens", 6120)    span.set_attribute("gen_ai.usage.output_tokens", 410)    span.set_attribute("gen_ai.response.finish_reasons", ["stop"])    span.set_attribute("app.prompt_version", "review-v9")    span.set_attribute("app.index_version", "repo-idx-2026-09-22")

This prints the span to the console; in production you swap in an OTLP exporter pointing at your collector. The span name follows the convention "operation model". In practice, auto-instrumentation libraries (for example OpenLLMetry or OpenInference) set the gen_ai.* attributes for you; you add the app.* ones.

Practical rules

  • Return the trace ID to the client and show it in support tools, so a complaint maps to a trace in one lookup.
  • Stamp the trace ID on every log line.
  • Store step inputs, not only the final answer, so you can replay a failing request against a fix.
  • Privacy: redact PII before storage; keep raw text for days or weeks, metrics for months; restrict access.

A real-life example

A code assistant for 2,000 engineers gets a report: "It suggested a function that does not exist." The engineer pastes the trace ID shown in the IDE plugin.

The trace shows: retrieval span 210 ms, returned 8 chunks from repo-idx-2026-09-22; rerank span kept 3; LLM span 3.9 s, 6,120 input tokens, finish_reason = stop. Opening the retrieved chunks shows they come from an old branch that was deleted last week — the index still holds it. The model did what it was told: it used the context. The fix is in the indexing job (drop deleted branches), not the prompt.

Finding this took four minutes. Before tracing, the team would have guessed "the model hallucinated" and edited the prompt.

Follow-up questions to expect

  • "Do you log 100% of prompts?" — Metrics and span metadata for 100%; full text for all or a sample depending on volume and privacy rules, always after redaction, with short retention.
  • "How do you trace an agent with many steps?" — Each loop iteration and tool call is a child span of one request trace, with step number and tool name as attributes.
  • "Why OpenTelemetry instead of a vendor SDK?" — One instrumentation, any backend; you can move from one observability tool to another without touching application code.