LLM Evaluation

Course Content

LLM Evaluation

6 sections · 50 lessons

Why are MCQ-based benchmarks popular, and what are their downsides?


What you need to know

Why they became the default

  • No ambiguity in grading.
  • Cheap: one short output, or even just the probability of each letter.
  • Easy to compare across models and over time.

Two ways to score

  • Generate the letter — ask the model to answer, extract "B" with a regex. Extraction failures count as wrong, so the regex matters.
  • Log-probability — compare the model's probability for each option. Works for base models; not how anyone uses a chat model.

These two methods give different numbers for the same model, which is one reason reported scores disagree.

Downsides

  1. Recognition, not generation. Picking the right option from four is easier than writing the right answer from scratch; the options give hints.
  2. Format sensitivity. Shuffling option order, changing "A)" to "(1)", or moving the correct option can change scores by several points. Some models prefer certain positions.
  3. Guessing floor. Random guessing gets 25% on four options. Rescale to see real ability above chance:
Text
above-chance accuracy = (raw - 0.25) / (1 - 0.25)raw 70%  ->  (0.70 - 0.25) / 0.75 = 60% of the way from guessing to perfect
  1. Contamination. MMLU (2020) questions are widely copied on the web.
  2. Label noise. A re-annotation effort, MMLU-Redux, found a noticeable share of MMLU questions with wrong keys or ambiguous wording, so no model can honestly reach 100%.

MMLU-Pro responded with ten options instead of four and harder, reasoning-heavy questions, lowering the guessing floor to 10%.

A real-life example

The bank's AI team is choosing a model for the complaint classifier. A vendor shows high scores on a banking-knowledge MCQ set. The team writes 50 MCQ questions from their own complaint categories, and both candidate models score above 90%.

Then they run the real task — reading a free-text Hinglish complaint and producing one of six labels with no options listed as hints in the text. The same models score 81% and 74%. The MCQ format had hidden the real difficulty: working out that "paisa kat gaya but merchant ko nahi mila" is a failed UPI payment, not fraud. The team decides on the real-task eval and keeps the MCQ set only as a quick knowledge check after fine-tuning.

Follow-up questions to expect

  • "How would you make an MCQ eval more robust?" — Shuffle option order across runs and average, use more options, report both generate-and-extract and log-prob methods consistently, and check for position bias.
  • "Are MCQ benchmarks useless?" — No; they are a cheap regression alarm. If a fine-tune drops MMLU sharply, it probably damaged general knowledge.
  • "What replaces them for real capability?" — Generative tasks with verifiable answers (maths, code with tests), agentic tasks with outcome checks, and your own task evals.