Course Content
LLM Evaluation
6 sections · 50 lessons
What is functional correctness, and how is pass@k used?
What you need to know
This is why code evaluation moved away from BLEU-style overlap: a function can differ from the reference in every line and still be right, or differ by one character (< versus <=) and be wrong.
pass@k
Drawing exactly k samples per problem gives a noisy estimate. The Codex paper (2021) draws n samples, counts the c that pass, and computes the chance that a random subset of k contains at least one pass:
pass@k = 1 - C(n - c, k) / C(n, k)C(a, b) is "a choose b". Averaged over problems. A numerically stable version:
1import numpy as np23def pass_at_k(n: int, c: int, k: int) -> float:4 """Unbiased pass@k: n samples drawn, c of them passed."""5 if n - c < k:6 return 1.07 return 1.0 - np.prod(1.0 - k / np.arange(n - c + 1, n + 1))89for k in (1, 3, 5):10 print(f"pass@{k} = {pass_at_k(n=10, c=3, k=k):.3f}")pass@1 = 0.300pass@3 = 0.708pass@5 = 0.917With 3 passing samples out of 10, one try succeeds 30% of the time, but if you can run five tries and a test suite picks the winner, you succeed 92% of the time. The check: pass@3 = 1 − C(7,3)/C(10,3) = 1 − 35/120 = 0.708.
How to read it
- pass@1 — what a user gets from one attempt without a verifier.
- pass@k for larger k — the ceiling when you can generate several candidates and verify them (tests, compiler, execution). It is not the same as reliability.
- pass^k (from agent benchmarks) — the chance that all k tries succeed; the opposite question.
Caveats
- Tests decide everything. Weak tests inflate the score; an empty function might pass a test that checks only the return type. SWE-bench Verified was created because many original tasks had unclear specs or unfair tests, and later audits found problems in it too.
- Sandbox execution. Run generated code without network access, with time and memory limits.
A real-life example
The code-review bot also suggests fixes. On 80 bug-fix tasks from the company's repositories, it generates 10 candidate patches per task, and each patch runs against the task's tests in a container. Averaged over tasks, pass@1 is 0.34 and pass@5 is 0.71.
That gap shapes the product: instead of posting one suggested fix, the bot generates five in parallel, runs the tests, and posts only a patch that passes — or says "no passing fix found". Posted fixes now pass tests in 71% of tasks, and no failing patch is ever shown. The cost is five generations plus five test runs per PR, which the team accepts for high-risk files only.
Follow-up questions to expect
- "Why not just generate k samples and check if any pass?" — That estimate has high variance; drawing n greater than k and using the formula gives an unbiased, lower-variance estimate from the same data.
- "What is the equivalent for SQL generation?" — Execution accuracy: run the generated and gold queries on the same database and compare result sets, ideally on several database states.
- "Does pass@k apply to non-code tasks?" — Wherever there is an automatic verifier: maths answers with a checker, extraction with a schema and gold values, tool calls with a sandbox.