LLM Evaluation

Course Content

LLM Evaluation

6 sections · 50 lessons

What is functional correctness, and how is pass@k used?


Ten sampled patches for one task, three passfailpassfailfailpassfailfailfailpassfail0123456789passpasspassn = 10, c = 3: pass@1 = 0.30, pass@3 = 0.71, pass@5 = 0.92.
pass@5 is only what users get if tests pick the winner; without a verifier they live with pass@1.

What you need to know

This is why code evaluation moved away from BLEU-style overlap: a function can differ from the reference in every line and still be right, or differ by one character (< versus <=) and be wrong.

pass@k

Drawing exactly k samples per problem gives a noisy estimate. The Codex paper (2021) draws n samples, counts the c that pass, and computes the chance that a random subset of k contains at least one pass:

Text
pass@k = 1 - C(n - c, k) / C(n, k)

C(a, b) is "a choose b". Averaged over problems. A numerically stable version:

Python
import numpy as npdef pass_at_k(n: int, c: int, k: int) -> float:    """Unbiased pass@k: n samples drawn, c of them passed."""    if n - c < k:        return 1.0    return 1.0 - np.prod(1.0 - k / np.arange(n - c + 1, n + 1))for k in (1, 3, 5):    print(f"pass@{k} = {pass_at_k(n=10, c=3, k=k):.3f}")
Text
pass@1 = 0.300pass@3 = 0.708pass@5 = 0.917

With 3 passing samples out of 10, one try succeeds 30% of the time, but if you can run five tries and a test suite picks the winner, you succeed 92% of the time. The check: pass@3 = 1 − C(7,3)/C(10,3) = 1 − 35/120 = 0.708.

How to read it

  • pass@1 — what a user gets from one attempt without a verifier.
  • pass@k for larger k — the ceiling when you can generate several candidates and verify them (tests, compiler, execution). It is not the same as reliability.
  • pass^k (from agent benchmarks) — the chance that all k tries succeed; the opposite question.

Caveats

  • Tests decide everything. Weak tests inflate the score; an empty function might pass a test that checks only the return type. SWE-bench Verified was created because many original tasks had unclear specs or unfair tests, and later audits found problems in it too.
  • Sandbox execution. Run generated code without network access, with time and memory limits.

A real-life example

The code-review bot also suggests fixes. On 80 bug-fix tasks from the company's repositories, it generates 10 candidate patches per task, and each patch runs against the task's tests in a container. Averaged over tasks, pass@1 is 0.34 and pass@5 is 0.71.

That gap shapes the product: instead of posting one suggested fix, the bot generates five in parallel, runs the tests, and posts only a patch that passes — or says "no passing fix found". Posted fixes now pass tests in 71% of tasks, and no failing patch is ever shown. The cost is five generations plus five test runs per PR, which the team accepts for high-risk files only.

Follow-up questions to expect

  • "Why not just generate k samples and check if any pass?" — That estimate has high variance; drawing n greater than k and using the formula gives an unbiased, lower-variance estimate from the same data.
  • "What is the equivalent for SQL generation?" — Execution accuracy: run the generated and gold queries on the same database and compare result sets, ideally on several database states.
  • "Does pass@k apply to non-code tasks?" — Wherever there is an automatic verifier: maths answers with a checker, extraction with a schema and gold values, tool calls with a sandbox.