Advanced
Open
Pro
Computing and Interpreting pass@k and pass^k for a Batch-Processing Agent
An agent processes end-of-day account reconciliations unattended, overnight, with no human reviewing each run. Testing shows a per-attempt success probability of p = 0.85 on the reconciliation task. The team currently reports only pass@5 to stakeholders, citing "99.99...% reliability."
- Compute pass@5 and pass^5 for this task, showing the arithmetic.
- Explain precisely why reporting only pass@5 is the wrong choice for this specific deployment context, using what the agent actually needs to do in production.
- If the team wants pass^5 to reach at least 90%, what per-attempt success probability p would be required (approximately), and what does that number tell you about how hard reliability improvement gets as k grows?
Share this question