Course Content
LLM Evaluation
6 sections · 50 lessons
What metrics should be tracked in production dashboards?
What you need to know
The four groups
| Group | Metrics |
|---|---|
| Reliability and performance | Volume; error rate by type (provider 5xx, timeout, rate limit, guardrail block); latency p50/p95/p99; time to first token; latency per stage |
| Cost | Input and output tokens; spend per request, user, feature; cache hit rate; cost per successful task; eval spend as a share of inference spend |
| Quality | Schema validity; required-field and citation presence; sampled faithfulness and relevance; empty-retrieval rate; refusal and over-refusal |
| User and business | Thumbs up/down; regenerate and rephrase rate; abandonment; escalation to human; task completion; the feature's business goal |
Cost per successful task
For an agent or multi-step flow, cost per call is misleading. If an agent costs ₹4 per run and succeeds 50% of the time, the cost per success is ₹8. A ₹6 agent that succeeds 90% of the time costs about ₹6.70 per success — cheaper where it matters.
Design rules
- Slice everything — by prompt version, model version, intent and user segment. A global average hides a regression in one intent.
- Deploy markers — most quality changes line up with a release, a data update or a provider change.
- Show uncertainty — daily judge scores from a sample have intervals; plot them so small moves are not over-read.
- One owner per chart — a metric nobody owns is not watched.
A real-life example
The bank's complaint-handling dashboard has one row per group. On a Monday, overall classification agreement with human review is flat at 91%. But the chart sliced by channel shows WhatsApp complaints dropping from 90% to 78% since Friday's deploy marker.
The traces show that Friday's release started sending voice-note transcripts from WhatsApp, which are long, informal and in mixed Hindi and English. The team adds a Hinglish slice to the golden set, routes voice transcripts through a clean-up step, and adds "complaint channel" as a permanent dashboard filter. Without slicing, the average would have hidden the problem for weeks, because WhatsApp is only 12% of volume.
Follow-up questions to expect
- "Which single metric would you watch for an LLM feature?" — The business outcome it exists for (resolution rate, task completion), backed by one quality metric and cost per success.
- "How do you choose alert thresholds?" — From historical variation: alert when a metric leaves its normal band for a sustained period, not on every blip.
- "Why track over-refusal?" — Because safety changes often make the system refuse normal requests; without the metric, users just quietly stop using it.