Advanced
Open
Pro
Best-of-N With a Verifier That Isn't as Good as You Think
Your team ships a math-tutoring feature using a small model (S) at N=9 verifier-scored best-of-N, matching the inference compute of one call to a larger model (B). Offline, S alone scores 34% pass@1; B alone scores 55% pass@1; and best-of-9 with your verifier was measured at 70% during the initial eval, so you shipped it. Three months later, a new harder problem category is added to the product, and pass rate on that category (measured after the fact) is only 31% — below even S's raw 34% pass@1 floor.
- Explain how a verifier-scored best-of-N result can land below the base model's own single-attempt accuracy. What does this imply about the verifier specifically on the new category?
- What measurement, run before shipping to the new category, would have caught this?
- Given this failure, would you (a) increase N, (b) switch to model B for this category, or (c) something else? Justify your choice.
Share this question