Course Content
LLM Evaluation
6 sections · 50 lessons
Why are MCQ-based benchmarks popular, and what are their downsides?
What you need to know
Why they became the default
- No ambiguity in grading.
- Cheap: one short output, or even just the probability of each letter.
- Easy to compare across models and over time.
Two ways to score
- Generate the letter — ask the model to answer, extract "B" with a regex. Extraction failures count as wrong, so the regex matters.
- Log-probability — compare the model's probability for each option. Works for base models; not how anyone uses a chat model.
These two methods give different numbers for the same model, which is one reason reported scores disagree.
Downsides
- Recognition, not generation. Picking the right option from four is easier than writing the right answer from scratch; the options give hints.
- Format sensitivity. Shuffling option order, changing "A)" to "(1)", or moving the correct option can change scores by several points. Some models prefer certain positions.
- Guessing floor. Random guessing gets 25% on four options. Rescale to see real ability above chance:
above-chance accuracy = (raw - 0.25) / (1 - 0.25)raw 70% -> (0.70 - 0.25) / 0.75 = 60% of the way from guessing to perfect- Contamination. MMLU (2020) questions are widely copied on the web.
- Label noise. A re-annotation effort, MMLU-Redux, found a noticeable share of MMLU questions with wrong keys or ambiguous wording, so no model can honestly reach 100%.
MMLU-Pro responded with ten options instead of four and harder, reasoning-heavy questions, lowering the guessing floor to 10%.
A real-life example
The bank's AI team is choosing a model for the complaint classifier. A vendor shows high scores on a banking-knowledge MCQ set. The team writes 50 MCQ questions from their own complaint categories, and both candidate models score above 90%.
Then they run the real task — reading a free-text Hinglish complaint and producing one of six labels with no options listed as hints in the text. The same models score 81% and 74%. The MCQ format had hidden the real difficulty: working out that "paisa kat gaya but merchant ko nahi mila" is a failed UPI payment, not fraud. The team decides on the real-task eval and keeps the MCQ set only as a quick knowledge check after fine-tuning.
Follow-up questions to expect
- "How would you make an MCQ eval more robust?" — Shuffle option order across runs and average, use more options, report both generate-and-extract and log-prob methods consistently, and check for position bias.
- "Are MCQ benchmarks useless?" — No; they are a cheap regression alarm. If a fine-tune drops MMLU sharply, it probably damaged general knowledge.
- "What replaces them for real capability?" — Generative tasks with verifiable answers (maths, code with tests), agentic tasks with outcome checks, and your own task evals.