LLMOps & Deployment

Course Content

LLMOps & Deployment

6 sections · 40 lessons

What does observability mean in the context of LLM systems?


What you need to know

Why normal API monitoring is not enough

In a normal API, a failure usually shows up as a 5xx error. In an LLM system, almost every bad outcome returns 200 OK: a hallucinated answer, a wrong tool call, an empty retrieval, a truncated reply. So you must capture the content — the prompt, the retrieved chunks and the response — or you cannot debug.

The three layers

  • Traces. One trace per user request. Each step is a span (a timed unit of work with attributes): embedding, retrieval, rerank, each LLM call, each tool call, each guardrail. The LLM span records model, prompt version, token counts, cost, latency and finish_reason.
  • Metrics. Aggregates over time: TTFT (time to first token), end-to-end p95, input and output tokens, cost per request and per tenant, cache hit rate, 429 and error rates, queue depth.
  • Quality signals. Schema-validation failures, refusal rate, guardrail blocks, groundedness scores, thumbs-down and regenerate rates, plus an LLM judge scoring 1–5% of traffic.

Standards and tools

OpenTelemetry (OTel) has GenAI semantic conventions — standard attribute names such as gen_ai.request.model and gen_ai.usage.input_tokens. They are still marked "Development" in 2026, but the core fields are settled and widely used. Tools that read this data include Langfuse and Arize Phoenix (both open source and self-hostable), LangSmith (from the LangChain team) and others. Using OTel means you can change tool without re-instrumenting.

Privacy

Prompts contain personal data. Redact PII before storing, keep raw text for a short time (for example 30 days), limit who can read it, and sample full payloads if volume is high. Metrics need no text and can be kept long.

A real-life example

A company runs an internal code assistant for 2,000 engineers. On Tuesday, the Slack channel fills with "it got worse today". The p95 latency and error-rate dashboards are green.

With observability, the on-call engineer filters traces by thumbs-down in the last day. Most bad traces share one attribute: index_version = repo-idx-2026-09-22. The retrieval spans show that the new index returns test files instead of source files for most queries, so the model writes code against the wrong functions. Rolling back the index fixes it in ten minutes. Without the index version on every span, the team would have spent days blaming the model.

Follow-up questions to expect

  • "What is the first thing you add to a new LLM app?" — Tracing with prompt version, model, tokens and cost on every LLM span; it is cheap and every later question needs it.
  • "How do you measure quality in production without labels?" — Proxies (schema failures, refusals, user feedback) plus an LLM judge on a sample, calibrated against human ratings.
  • "Do you log every prompt?" — Metrics for all requests, full text for a sample or with short retention, always after PII redaction.