LLMOps & Deployment

Course Content

LLMOps & Deployment

6 sections · 40 lessons

How do you monitor LLM applications in real time?


What you need to know

The three tiers

TierMetricsAlert example
System healthRPS, error rate by type, TTFT and e2e p95/p99, queue depth, KV-cache usage, GPU healthQueue depth over 50 for 2 minutes
Cost and usageInput/output tokens per request, spend per hour, per tenant, per feature; cache hit rateSpend rate 1.5× the same hour last week
Quality proxiesSchema-failure rate, guardrail blocks, refusals, finish_reason = length, empty retrievals, thumbs-down, escalations, sampled judge scoreThumbs-down rate doubles in an hour

Why each tier matters

  • System health catches outages and overload. Most of it is standard SRE work.
  • Cost catches bugs that nothing else catches. A prompt change that doubles the context length is invisible to latency alerts at first, but the spend rate jumps within minutes.
  • Quality proxies catch the LLM-specific failures that return 200 OK. No single proxy is quality, but together they move when quality moves.

Leading and lagging signals

Latency is a lagging signal: when p95 breaks, users are already waiting. Queue depth and KV-cache usage are leading signals: they rise before latency does. Page on leading signals; use latency for SLO reporting.

Alert rules, in Prometheus form

YAML
groups:  - name: support-bot    rules:      - alert: QueueBuilding        expr: sum(vllm:num_requests_waiting) > 50        for: 2m        labels: {severity: page}      - alert: TTFTHigh        expr: |          histogram_quantile(0.95,            sum(rate(vllm:time_to_first_token_seconds_bucket[5m])) by (le)) > 1.0        for: 5m        labels: {severity: page}      - alert: SpendRateHigh        expr: sum(rate(llm_cost_usd_total[15m])) > 1.5 * sum(rate(llm_cost_usd_total[15m] offset 1w))        for: 15m        labels: {severity: ticket}

The first two use metrics that vLLM exports. llm_cost_usd_total is a counter your gateway would emit. The third rule compares with the same time last week (offset 1w), because traffic has a daily and weekly shape and a fixed threshold would fire every evening.

Online quality sampling

Send 1–5% of traffic to an LLM judge that scores groundedness or helpfulness, and track the rolling average. Calibrate the judge against human ratings first, and pin the judge's model and prompt.

A real-life example

A fintech support bot on the first evening of its Diwali sale:

  • 7:40 pm — SpendRateHigh fires: spend is 2.3× last week's rate while traffic is only 1.6×. The cost dashboard, sliced by prompt version, shows a new promotion prompt pasting the full 3,000-token sale terms into every call. It is moved into a cached prefix; spend per request falls back.
  • 8:10 pm — QueueBuilding fires. TTFT p95 is still 0.8 s, but the autoscaler is already adding replicas; by the time users arrive in full, capacity is there.
  • 9:30 pm — the quality dashboard shows escalations to human agents up from 9% to 15% for refund questions only. Sampled traces show the bot does not know the sale's new refund window. The knowledge base is updated at 10 pm.

None of these three issues would show up on a server-health dashboard alone.

Follow-up questions to expect

  • "Which single metric would you page on?" — Queue depth (or queue wait time), because it leads latency and directly measures unmet demand.
  • "How do you measure quality without labels?" — Proxies plus a calibrated LLM judge on a sample; review a few low-scoring traces by hand every week.
  • "How do you avoid alert fatigue?" — Page only on user-facing and leading signals, compare with last week rather than fixed thresholds, and send cost and quality drifts to tickets.