Course Content
LLM Evaluation
6 sections · 50 lessons
How do cost and latency impact the use of AI judges in production?
What you need to know
In the request path or not
- Online gate (synchronous) — the judge runs before the user sees the answer and can block it. Adds latency to every request. Use only for high-risk checks, and prefer a small, fast classifier.
- Background scoring (asynchronous) — the answer is returned; the trace is logged; a worker scores it later. No user-facing latency. This is how LangSmith and Langfuse online evaluators work.
Sampling
- A random 1 to 5% of traffic for trend lines.
- 100% of flagged traces: thumbs-down, regenerations, errors, escalations, high-risk intents.
- Stratify by intent so low-volume intents still get enough samples.
Cost arithmetic
Suppose the bank's reply drafter handles 200,000 drafts a day. A judge call reads about 1,500 tokens (complaint, draft, rubric) and writes about 150.
judge everything: 200,000 judge calls/day3% sample + flagged: 6,000 + ~2,000 flagged = ~8,000 calls/day (about 25x fewer)At ~8,000 calls a day, a 3% sample still gives a daily faithfulness estimate with an interval of about ±1 point, which is enough to see a real regression within a day.
Other savings
- Deterministic checks first — schema, regex, citation presence — then judge only what passes.
- Small judges for binary checks, strong judges for the uncertain slice.
- Caching on a hash of (input, output, judge prompt, judge version).
- Batch APIs for offline suites where results can wait hours, often at a lower price.
- Short outputs — binary verdicts with brief reasons cost less than long critiques.
A useful budget number: eval spend as a percentage of inference spend. Teams often aim for a small fraction such as 5 to 15%; well above that usually means over-judging.
A real-life example
The HR assistant first ran three judge checks synchronously on every answer — faithfulness, relevance and tone — to "guarantee quality". p95 latency went from 3.1 to 7.8 seconds, and the monthly model bill nearly tripled.
The redesign: a small, fast safety classifier stays in the request path (it adds about 150 ms). Faithfulness and relevance move to a background worker scoring 5% of conversations plus every thumbs-down. Tone is checked weekly on 200 samples. Latency drops back to 3.3 seconds, eval spend falls to about 8% of inference spend, and the dashboard still shows faithfulness daily.
Follow-up questions to expect
- "When must a judge be synchronous?" — When a bad output causes harm the moment it is shown: unsafe content, leaked personal data, or a regulated claim. Use the cheapest reliable check that does the job.
- "How big should the sample be?" — Big enough that the daily or weekly interval is smaller than the change you want to detect; compute it from the expected rate.
- "Can you fine-tune a small judge?" — Yes; train a small classifier on the strong judge's calibrated labels plus human labels for high-volume checks, and keep checking it against humans.