LLM Evaluation

Course Content

LLM Evaluation

6 sections · 50 lessons

How do offline and online evaluations differ in practice?


What you need to know

OfflineOnline
Ground truthCurated labels or rubricsUser behaviour, mostly implicit
SpeedMinutesDays to weeks
RiskNoneReal users see the change
RepeatableYes, same inputs every runNo, traffic changes daily
Blind spotOnly tests what you thought ofNoisy, many causes for any change

Online signals

  • Explicit feedback — thumbs up/down, ratings. Honest but sparse; often under 1% of users click.
  • Implicit feedback — copying the answer, regenerating, rephrasing the same question, abandoning the chat, escalating to a human agent. High volume, but you must decide what each means.
  • Sampled judge scores — an LLM judge scores a few percent of production traces for faithfulness or safety, in the background.
  • Business metrics in an A/B test — conversion, resolution rate, handle time, compared between a control and a treatment group.

Why you need both

They fail in opposite directions. Offline scores can rise while users get worse, because the eval set no longer looks like real traffic. Online metrics can move for reasons unrelated to your change — a festival sale, a new user segment, an outage. Offline answers "is it better on what I know?"; online answers "did it matter?".

The feedback loop

Every interesting production failure — a thumbs-down, an escalation, a complaint — is reviewed and, if it is a real failure, added as an offline test case. Over months, this keeps the offline set close to real usage.

A real-life example

The e-commerce description generator gets a new prompt. Offline, on 300 stored spec sheets, a pairwise LLM judge prefers the new descriptions 64% of the time, and the deterministic checks still pass at 100%. The team ships it to 10% of product pages as an online A/B test for two weeks.

Result: add-to-cart rate is unchanged within its confidence interval, but "item not as described" returns fall slightly on the treatment pages. The judge had rewarded livelier writing; users cared about accuracy. The team keeps the new prompt for the accuracy gain, and adds five return complaints as new offline cases about size and colour claims.

Follow-up questions to expect

  • "Which do you run first?" — Offline, always. It is cheap and protects users; online testing is only for candidates that already pass the offline gate.
  • "How do you get labels online?" — Mostly you do not. You rely on reference-free checks, sampled judge scores, user behaviour, and a human review queue for flagged traces.
  • "What if offline and online disagree?" — Trust the online result for the product decision, then investigate why the offline set missed it; usually the dataset has drifted from real traffic.