Course Content
LLM Evaluation
6 sections · 50 lessons
How do you monitor models continuously after deployment?
What you need to know
Layer 1: telemetry
Tools such as LangSmith, Langfuse and Arize Phoenix collect traces; OpenTelemetry's generative-AI semantic conventions define standard attribute names so traces can flow into general observability tools. Without traces, you can detect a problem but not debug it.
Layer 2: automatic quality signals
- 100% of traffic, cheap checks — schema validity, refusal detection, empty retrieval, guardrail hits, length anomalies, forbidden phrases.
- Sampled judge scores — 1 to 5% of traffic plus all flagged traces: faithfulness, answer relevance, safety.
Layer 3: user and human signals
- Explicit: thumbs up/down, ratings.
- Implicit: regenerate, rephrase, copy, abandon, escalate to a human.
- Human review: a weekly stratified sample graded by a domain expert — the check on everything else.
Alerting
- Threshold alerts for hard failures: error rate, p95 latency, cost per request.
- Distribution alerts for quality: a shift in topic mix, retrieval score distribution, refusal rate or judge score distribution compared with the last few weeks.
- Canary set — 50 fixed prompts with known good outputs run every hour or day. A change in their outputs with no deploy means something changed upstream.
A real-life example
The e-commerce description generator writes 5,000 descriptions a day. One Tuesday, nothing errors, latency is normal, and cost is flat. But the deterministic check "every spec attribute mentioned" drops from 99% to 93%, and the hourly canary shows 6 of 50 descriptions now missing the warranty line.
The traces show why: a catalogue team changed the spec field name from warranty to warranty_period, and the prompt template rendered an empty value. No model or prompt changed. The fix took ten minutes; without the attribute check and canary, it would have surfaced weeks later as customer questions about missing warranties.
Follow-up questions to expect
- "What would you alert on first?" — Error rate, latency, cost, and the deterministic quality checks; they are cheap, precise and catch the most common breakages.
- "How do you catch a provider changing the model?" — Pin model versions, and run the canary set on a schedule; alert when canary outputs or scores change without a deploy.
- "How do you handle privacy in traces?" — Redact or mask personal data before storage, restrict access, and set retention limits.