Course Content
LLMOps & Deployment
6 sections · 40 lessons
How do you monitor LLM applications in real time?
What you need to know
The three tiers
| Tier | Metrics | Alert example |
|---|---|---|
| System health | RPS, error rate by type, TTFT and e2e p95/p99, queue depth, KV-cache usage, GPU health | Queue depth over 50 for 2 minutes |
| Cost and usage | Input/output tokens per request, spend per hour, per tenant, per feature; cache hit rate | Spend rate 1.5× the same hour last week |
| Quality proxies | Schema-failure rate, guardrail blocks, refusals, finish_reason = length, empty retrievals, thumbs-down, escalations, sampled judge score | Thumbs-down rate doubles in an hour |
Why each tier matters
- System health catches outages and overload. Most of it is standard SRE work.
- Cost catches bugs that nothing else catches. A prompt change that doubles the context length is invisible to latency alerts at first, but the spend rate jumps within minutes.
- Quality proxies catch the LLM-specific failures that return 200 OK. No single proxy is quality, but together they move when quality moves.
Leading and lagging signals
Latency is a lagging signal: when p95 breaks, users are already waiting. Queue depth and KV-cache usage are leading signals: they rise before latency does. Page on leading signals; use latency for SLO reporting.
Alert rules, in Prometheus form
1groups:2 - name: support-bot3 rules:4 - alert: QueueBuilding5 expr: sum(vllm:num_requests_waiting) > 506 for: 2m7 labels: {severity: page}8 - alert: TTFTHigh9 expr: |10 histogram_quantile(0.95,11 sum(rate(vllm:time_to_first_token_seconds_bucket[5m])) by (le)) > 1.012 for: 5m13 labels: {severity: page}14 - alert: SpendRateHigh15 expr: sum(rate(llm_cost_usd_total[15m])) > 1.5 * sum(rate(llm_cost_usd_total[15m] offset 1w))16 for: 15m17 labels: {severity: ticket}The first two use metrics that vLLM exports. llm_cost_usd_total is a counter your gateway would emit. The third rule compares with the same time last week (offset 1w), because traffic has a daily and weekly shape and a fixed threshold would fire every evening.
Online quality sampling
Send 1–5% of traffic to an LLM judge that scores groundedness or helpfulness, and track the rolling average. Calibrate the judge against human ratings first, and pin the judge's model and prompt.
A real-life example
A fintech support bot on the first evening of its Diwali sale:
- 7:40 pm —
SpendRateHighfires: spend is 2.3× last week's rate while traffic is only 1.6×. The cost dashboard, sliced by prompt version, shows a new promotion prompt pasting the full 3,000-token sale terms into every call. It is moved into a cached prefix; spend per request falls back. - 8:10 pm —
QueueBuildingfires. TTFT p95 is still 0.8 s, but the autoscaler is already adding replicas; by the time users arrive in full, capacity is there. - 9:30 pm — the quality dashboard shows escalations to human agents up from 9% to 15% for refund questions only. Sampled traces show the bot does not know the sale's new refund window. The knowledge base is updated at 10 pm.
None of these three issues would show up on a server-health dashboard alone.
Follow-up questions to expect
- "Which single metric would you page on?" — Queue depth (or queue wait time), because it leads latency and directly measures unmet demand.
- "How do you measure quality without labels?" — Proxies plus a calibrated LLM judge on a sample; review a few low-scoring traces by hand every week.
- "How do you avoid alert fatigue?" — Page only on user-facing and leading signals, compare with last week rather than fixed thresholds, and send cost and quality drifts to tickets.