LLM Evaluation

Course Content

LLM Evaluation

6 sections · 50 lessons

How do you measure similarity between two open-ended texts?


What you need to know

"Similar" can mean same words, same topic, or same meaning. Each level has its own tools.

Level 1: lexical

Token F1, Jaccard overlap of n-grams, ROUGE-L, edit distance. Fast and explainable. Good for near-duplicate detection and for spotting when a model copies a template. Useless for paraphrase: "The package arrives Tuesday" and "Delivery is expected on Tuesday" share little.

Level 2: embeddings

Embeddings capture paraphrase, but they measure closeness of topic and phrasing, not truth. Sentences differing by one number, a "not", or swapped roles ("the bank refunded the customer" versus "the customer refunded the bank") usually get very high similarity. Also, raw cosine values have no universal meaning: 0.8 can mean "same" for one model and "loosely related" for another. Calibrate a threshold on 100 to 200 labelled pairs from your data.

Level 3: entailment

Ask whether A implies B and whether B implies A, with an NLI model or an LLM judge. If both directions hold, they mean the same. If A implies B but not the reverse, A contains extra information. Entailment catches negation and number changes that embeddings miss.

Claim decomposition

For factual texts, split each into atomic claims — one fact each — and match them:

Text
Gold:   leave is 15 working days | applies after probation | needs manager approvalAnswer: leave is 15 working days | needs manager approval | can be split into two partsclaim recall    = 2/3 gold claims presentclaim precision = 2/3 answer claims supported by gold (the third is new, maybe invented)

This gives an interpretable score and names the missing fact ("applies after probation").

A real-life example

The bank wants to find duplicate complaints: the same customer often writes to email, the app and Twitter about one failed UPI payment. Lexical matching missed most duplicates because customers reword each message. Embedding similarity with a 0.85 threshold found them, but also merged different complaints about the same merchant with different amounts.

The final design uses embeddings to find candidates (cheap, high recall), then a deterministic check that amount, date and transaction reference match, and an entailment check for the rest. On 500 labelled pairs, precision rose from 0.74 with embeddings alone to 0.96 with the combined approach, at about the same recall.

Follow-up questions to expect

  • "Why not just use BERTScore?" — It is token-level embedding similarity, so it inherits the same weakness with numbers and negation; fine for paraphrase tolerance, not for fact checking.
  • "How do you choose a cosine threshold?" — Label pairs as same/different, plot precision and recall at different thresholds, and choose based on the cost of each error.
  • "How reliable is claim decomposition?" — It depends on the decomposer; LLMs sometimes split or merge claims inconsistently, so fix the prompt, use examples, and spot-check.