LLM Evaluation

Course Content

LLM Evaluation

6 sections · 50 lessons

How do cost and latency impact the use of AI judges in production?


What you need to know

In the request path or not

  • Online gate (synchronous) — the judge runs before the user sees the answer and can block it. Adds latency to every request. Use only for high-risk checks, and prefer a small, fast classifier.
  • Background scoring (asynchronous) — the answer is returned; the trace is logged; a worker scores it later. No user-facing latency. This is how LangSmith and Langfuse online evaluators work.

Sampling

  • A random 1 to 5% of traffic for trend lines.
  • 100% of flagged traces: thumbs-down, regenerations, errors, escalations, high-risk intents.
  • Stratify by intent so low-volume intents still get enough samples.

Cost arithmetic

Suppose the bank's reply drafter handles 200,000 drafts a day. A judge call reads about 1,500 tokens (complaint, draft, rubric) and writes about 150.

Text
judge everything:      200,000 judge calls/day3% sample + flagged:   6,000 + ~2,000 flagged = ~8,000 calls/day  (about 25x fewer)

At ~8,000 calls a day, a 3% sample still gives a daily faithfulness estimate with an interval of about ±1 point, which is enough to see a real regression within a day.

Other savings

  • Deterministic checks first — schema, regex, citation presence — then judge only what passes.
  • Small judges for binary checks, strong judges for the uncertain slice.
  • Caching on a hash of (input, output, judge prompt, judge version).
  • Batch APIs for offline suites where results can wait hours, often at a lower price.
  • Short outputs — binary verdicts with brief reasons cost less than long critiques.

A useful budget number: eval spend as a percentage of inference spend. Teams often aim for a small fraction such as 5 to 15%; well above that usually means over-judging.

A real-life example

The HR assistant first ran three judge checks synchronously on every answer — faithfulness, relevance and tone — to "guarantee quality". p95 latency went from 3.1 to 7.8 seconds, and the monthly model bill nearly tripled.

The redesign: a small, fast safety classifier stays in the request path (it adds about 150 ms). Faithfulness and relevance move to a background worker scoring 5% of conversations plus every thumbs-down. Tone is checked weekly on 200 samples. Latency drops back to 3.3 seconds, eval spend falls to about 8% of inference spend, and the dashboard still shows faithfulness daily.

Follow-up questions to expect

  • "When must a judge be synchronous?" — When a bad output causes harm the moment it is shown: unsafe content, leaked personal data, or a regulated claim. Use the cheapest reliable check that does the job.
  • "How big should the sample be?" — Big enough that the daily or weekly interval is smaller than the change you want to detect; compute it from the expected rate.
  • "Can you fine-tune a small judge?" — Yes; train a small classifier on the strong judge's calibrated labels plus human labels for high-volume checks, and keep checking it against humans.