LLM Evaluation

Course Content

LLM Evaluation

6 sections · 50 lessons

How do you monitor models continuously after deployment?


What you need to know

Layer 1: telemetry

Tools such as LangSmith, Langfuse and Arize Phoenix collect traces; OpenTelemetry's generative-AI semantic conventions define standard attribute names so traces can flow into general observability tools. Without traces, you can detect a problem but not debug it.

Layer 2: automatic quality signals

  • 100% of traffic, cheap checks — schema validity, refusal detection, empty retrieval, guardrail hits, length anomalies, forbidden phrases.
  • Sampled judge scores — 1 to 5% of traffic plus all flagged traces: faithfulness, answer relevance, safety.

Layer 3: user and human signals

  • Explicit: thumbs up/down, ratings.
  • Implicit: regenerate, rephrase, copy, abandon, escalate to a human.
  • Human review: a weekly stratified sample graded by a domain expert — the check on everything else.

Alerting

  • Threshold alerts for hard failures: error rate, p95 latency, cost per request.
  • Distribution alerts for quality: a shift in topic mix, retrieval score distribution, refusal rate or judge score distribution compared with the last few weeks.
  • Canary set — 50 fixed prompts with known good outputs run every hour or day. A change in their outputs with no deploy means something changed upstream.

A real-life example

The e-commerce description generator writes 5,000 descriptions a day. One Tuesday, nothing errors, latency is normal, and cost is flat. But the deterministic check "every spec attribute mentioned" drops from 99% to 93%, and the hourly canary shows 6 of 50 descriptions now missing the warranty line.

The traces show why: a catalogue team changed the spec field name from warranty to warranty_period, and the prompt template rendered an empty value. No model or prompt changed. The fix took ten minutes; without the attribute check and canary, it would have surfaced weeks later as customer questions about missing warranties.

Follow-up questions to expect

  • "What would you alert on first?" — Error rate, latency, cost, and the deterministic quality checks; they are cheap, precise and catch the most common breakages.
  • "How do you catch a provider changing the model?" — Pin model versions, and run the canary set on a schedule; alert when canary outputs or scores change without a deploy.
  • "How do you handle privacy in traces?" — Redact or mask personal data before storage, restrict access, and set retention limits.