LLMOps & Deployment

Course Content

LLMOps & Deployment

6 sections · 40 lessons

What tools and metrics are needed for effective observability?


What you need to know

Tools by layer (2026)

LayerToolsNotes
Instrumentation standardOpenTelemetry + GenAI semantic conventions; OpenLLMetry, OpenInference auto-instrumentationInstrument once, send anywhere
LLM tracing, prompts, evalsLangfuse (open source), Arize Phoenix (open source), LangSmith, Braintrust, HeliconeTrace view with prompts, datasets, judges
Metrics and dashboardsPrometheus + Grafana, or your APMTime series and alerts
Serving metricsvLLM, SGLang /metrics endpointsTTFT, ITL, queue, KV cache
GPU fleetNVIDIA DCGM exporterUtilisation, memory, temperature, ECC errors, throttling
Cost accountingGateway (for example LiteLLM-style proxies)Tokens and spend per key, team, model
Offline evalspromptfoo, DeepEval, RagasCI eval gates, RAG metrics

Useful vLLM metrics

Names change between versions, so check your version's docs. In current vLLM:

  • vllm:time_to_first_token_seconds and vllm:inter_token_latency_seconds (histograms)
  • vllm:e2e_request_latency_seconds
  • vllm:num_requests_running and vllm:num_requests_waiting
  • vllm:kv_cache_usage_perc (older versions used vllm:gpu_cache_usage_perc)
  • vllm:num_preemptions_total — requests evicted because the KV cache was full
  • vllm:prompt_tokens_total and vllm:generation_tokens_total

The metrics that earn their place

  • Latency: TTFT, TPOT, end-to-end p95/p99
  • Throughput: output tokens/s per GPU
  • Saturation: queue depth, KV-cache usage, preemptions
  • Cost: tokens and dollars per request, per tenant, per feature; cache hit rate
  • Reliability: error rate by type, 429 rate, fallback rate
  • Quality: schema-failure, guardrail-block, refusal and groundedness rates; retrieval hit rate; thumbs-down and escalations

The rule that ties it together

Every metric carries the same labels: prompt version, model snapshot, route, tenant. A dashboard that shows a spike but cannot say which variant caused it does not shorten an incident.

A real-life example

A hospital's summariser must stay on-prem, so SaaS tracing tools are ruled out. The team builds its stack from self-hosted pieces:

  • Langfuse, self-hosted inside the hospital network, receives OTel traces with prompts and summaries. Access is limited to three engineers and the clinical safety officer; raw text is deleted after 30 days.
  • Prometheus and Grafana scrape vLLM and DCGM. A dashboard shows TTFT, queue depth, KV-cache usage and GPU temperature per server.
  • A nightly job runs the 400-case eval set and writes the score to Prometheus as a gauge, so quality sits on the same dashboard as latency.

When a GPU starts throttling on hot afternoons (DCGM shows clocks dropping), the latency panel explains itself, and facilities fix the server-room cooling.

Follow-up questions to expect

  • "Langfuse, LangSmith or Phoenix?" — All trace, evaluate and manage datasets. Langfuse and Phoenix are open source and easy to self-host; LangSmith fits teams already on LangChain/LangGraph. Choose by hosting rules and ecosystem, and keep OTel as the standard so you can move.
  • "What would you build first?" — Tracing with tokens and cost on every call, then a cost-per-tenant dashboard, then quality proxies.
  • "How do you get quality metrics onto the same dashboard?" — Emit judge scores and eval results as metrics with the same labels as latency and cost.