Glossary ¡ 5 minute read
What Is Pass@k? Measuring AI Success Across Attempts
Pass@k measures the probability that at least one of k generated attempts is correct. It reflects real value only when something can identify which attempt succeeded â a test suite, a validator, a human. Without a verifier, a high pass@k means the right answer was produced and not recognised.
Pass@k is a useful metric in the setting it was designed for and a misleading one almost everywhere else. Its popularity in benchmark reporting means it frequently appears in capability claims about systems that have no way to exploit it, which inflates expectations that production then fails to meet. This explainer covers when it means something. It complements what is an evaluation rubric and ai evaluation checklist, and reflects FISTA Solutions' approach in AI enablement delivery.
What does pass@k measure?
The probability that at least one of k independently generated attempts is correct. Pass@1 is the chance the first attempt works; pass@10 is the chance that any of ten does.
The gap between them is often large, which is what makes the metric attractive to report and dangerous to interpret without qualification.
| Setting | Verifier available | Pass@k meaningful |
|---|---|---|
| Code with tests | Test suite | Yes |
| Structured output | Schema validator | Yes |
| Calculation | Deterministic recomputation | Yes |
| SQL generation | Execution and constraint check | Partly |
| Summarisation | None | No |
| Advice or explanation | None | No |
Why does it work for code?
Because tests verify. Generating ten candidate functions and running the test suite against each identifies which one is correct, automatically and cheaply. The workflow the metric describes is a workflow that actually exists.
That is the whole condition. Pass@k is meaningful exactly where something can select the successful attempt without a human examining all of them.
Where does it mislead?
Everywhere a verifier is absent. A summarisation system reporting pass@10 produced one acceptable summary among ten and cannot tell which. The user receives the first one generated, so the experienced quality is pass@1.
Reporting the higher number describes a capability that exists in the model and cannot be delivered to anyone, which is a meaningful distinction that the metric's presentation tends to obscure.
What qualifies as a verifier?
Something that distinguishes correct from incorrect reliably and without human reading. Test suites, schema validators, compilers, database constraints, and deterministic recalculation all qualify.
Model-based scoring is weaker: a judge that is itself sometimes wrong turns selection into a probabilistic improvement rather than a guarantee, and its reliability must be measured on the specific task before it can be treated as a verifier at all. See what is an evaluation rubric.
What does k cost?
Roughly linearly. Ten attempts cost about ten times the tokens of one, plus verification. For a high-value low-volume task â generating a migration script, drafting a complex query â that is easily justified. For a high-volume interactive task it usually is not.
Best-of-n sampling is the production form of this trade, and it should be applied selectively rather than as a default setting.
What should production systems report?
Pass@1 where the first output is what users receive. Pass@k with the k stated, and the verifier described, where selection genuinely happens. An unqualified pass rate is uninterpretable, and comparing a pass@10 from one system against a pass@1 from another is not a comparison at all.
How should vendor claims be read?
By asking what k was used and what verified the result. A benchmark figure quoted without those is not evidence about what the system will do in your application, and the difference between the quoted number and what you would experience can be very large.
What should you do first?
For each of your AI tasks, ask whether a verifier exists or could be built cheaply. Where one can, best-of-n becomes available as a quality lever. Where none can, pass@1 is your real performance and any higher number is describing something you cannot deliver.
How does this relate to agent evaluation?
Directly, because agents retry. An agent that attempts a task, observes failure, and tries a different approach is performing verification-driven selection inside its own loop, and its effective performance resembles pass@k where k is the iteration limit.
That makes the verifier question central to agent design rather than only to metrics. An agent that can tell whether its action succeeded improves with more iterations; one that cannot merely spends more to produce the same uncertainty.
What about human verification?
It counts, and it changes the economics rather than the logic. Where a person reviews output anyway, generating several candidates and presenting the best few can improve outcomes at acceptable cost, because the human is the verifier the metric requires.
The design detail that matters is how the candidates are presented. Showing three options invites comparison and improves the result; showing one and quietly discarding two wastes the generation entirely.
How should benchmarks be compared?
Only at equal k, with the same verifier, on the same tasks. Those three conditions are rarely all met in published comparisons, which is why benchmark tables should be read as a rough ordering rather than as measurements of the difference between systems.
How FISTA Solutions helps
FISTA Solutions identifies where verifiers exist or can be built, applies best-of-n selectively to high-value tasks, reports pass@1 for systems without verification, validates model-based judges before treating them as selectors, and reads vendor benchmark claims against the k and the verifier used, through AI enablement, AI agents, and forward deployed engineers. The record behind the approach is 150+ projects for 50+ companies with 99.9% uptime.
To measure AI capability in terms your users will actually experience, message FISTA on WhatsApp, or read what is an evaluation rubric.
Share-ready article cover
Download the generated social format.
Clear answers
Questions raised by this field note.
Straightforward guidance for evaluating scope, fit, and the next step.
01Where does pass@k come from?
Code generation benchmarks, where it fits naturally because a test suite can determine which attempt is correct. Generating several candidates and keeping the one that passes is a real workflow, so the metric describes something achievable.
02When does it mislead?
When no verifier exists. A summarisation system with a high pass@10 produced one good summary among ten and has no way to identify it, so the user still receives whichever one was generated first. The metric describes a capability nobody can use.
03What counts as a verifier?
Anything that reliably distinguishes correct from incorrect without a human reading everything: a test suite, a schema validator, a compiler, a database constraint, a deterministic calculation. Model-based scoring is weaker and needs its own validation before it counts.
04Why does k matter for cost?
Because each attempt is a separate generation. Pass@10 costs roughly ten times pass@1 in tokens, plus the cost of verification. That trade can be worthwhile for high-value low-volume tasks and is rarely worthwhile for high-volume interactive ones where latency also matters.
05What should production systems report?
Pass@1 where the first output is what users see, and pass@k with k stated where a verifier selects among attempts. An unqualified pass rate without the k is uninterpretable, and comparisons across different k values are meaningless.
Continue exploring
Related capabilities
Start with the hard problem
Need the outcome owned, not merely analyzed?
Tell us where delivery is constrained. Weâll map the fastest credible path from intent to verified production.