Course Content
LLMOps & Deployment
6 sections · 40 lessons
What tools and metrics are needed for effective observability?
What you need to know
Tools by layer (2026)
| Layer | Tools | Notes |
|---|---|---|
| Instrumentation standard | OpenTelemetry + GenAI semantic conventions; OpenLLMetry, OpenInference auto-instrumentation | Instrument once, send anywhere |
| LLM tracing, prompts, evals | Langfuse (open source), Arize Phoenix (open source), LangSmith, Braintrust, Helicone | Trace view with prompts, datasets, judges |
| Metrics and dashboards | Prometheus + Grafana, or your APM | Time series and alerts |
| Serving metrics | vLLM, SGLang /metrics endpoints | TTFT, ITL, queue, KV cache |
| GPU fleet | NVIDIA DCGM exporter | Utilisation, memory, temperature, ECC errors, throttling |
| Cost accounting | Gateway (for example LiteLLM-style proxies) | Tokens and spend per key, team, model |
| Offline evals | promptfoo, DeepEval, Ragas | CI eval gates, RAG metrics |
Useful vLLM metrics
Names change between versions, so check your version's docs. In current vLLM:
vllm:time_to_first_token_secondsandvllm:inter_token_latency_seconds(histograms)vllm:e2e_request_latency_secondsvllm:num_requests_runningandvllm:num_requests_waitingvllm:kv_cache_usage_perc(older versions usedvllm:gpu_cache_usage_perc)vllm:num_preemptions_total— requests evicted because the KV cache was fullvllm:prompt_tokens_totalandvllm:generation_tokens_total
The metrics that earn their place
- Latency: TTFT, TPOT, end-to-end p95/p99
- Throughput: output tokens/s per GPU
- Saturation: queue depth, KV-cache usage, preemptions
- Cost: tokens and dollars per request, per tenant, per feature; cache hit rate
- Reliability: error rate by type, 429 rate, fallback rate
- Quality: schema-failure, guardrail-block, refusal and groundedness rates; retrieval hit rate; thumbs-down and escalations
The rule that ties it together
Every metric carries the same labels: prompt version, model snapshot, route, tenant. A dashboard that shows a spike but cannot say which variant caused it does not shorten an incident.
A real-life example
A hospital's summariser must stay on-prem, so SaaS tracing tools are ruled out. The team builds its stack from self-hosted pieces:
- Langfuse, self-hosted inside the hospital network, receives OTel traces with prompts and summaries. Access is limited to three engineers and the clinical safety officer; raw text is deleted after 30 days.
- Prometheus and Grafana scrape vLLM and DCGM. A dashboard shows TTFT, queue depth, KV-cache usage and GPU temperature per server.
- A nightly job runs the 400-case eval set and writes the score to Prometheus as a gauge, so quality sits on the same dashboard as latency.
When a GPU starts throttling on hot afternoons (DCGM shows clocks dropping), the latency panel explains itself, and facilities fix the server-room cooling.
Follow-up questions to expect
- "Langfuse, LangSmith or Phoenix?" — All trace, evaluate and manage datasets. Langfuse and Phoenix are open source and easy to self-host; LangSmith fits teams already on LangChain/LangGraph. Choose by hosting rules and ecosystem, and keep OTel as the standard so you can move.
- "What would you build first?" — Tracing with tokens and cost on every call, then a cost-per-tenant dashboard, then quality proxies.
- "How do you get quality metrics onto the same dashboard?" — Emit judge scores and eval results as metrics with the same labels as latency and cost.