Match a job Paths Subjects Questions Quizzes Pricing
Advanced Open Pro

Best-of-N With a Verifier That Isn't as Good as You Think

Your team ships a math-tutoring feature using a small model (S) at N=9 verifier-scored best-of-N, matching the inference compute of one call to a larger model (B). Offline, S alone scores 34% pass@1; B alone scores 55% pass@1; and best-of-9 with your verifier was measured at 70% during the initial eval, so you shipped it. Three months later, a new harder problem category is added to the product, and pass rate on that category (measured after the fact) is only 31% — below even S's raw 34% pass@1 floor.

  1. Explain how a verifier-scored best-of-N result can land below the base model's own single-attempt accuracy. What does this imply about the verifier specifically on the new category?
  2. What measurement, run before shipping to the new category, would have caught this?
  3. Given this failure, would you (a) increase N, (b) switch to model B for this category, or (c) something else? Justify your choice.

Share this question

← Back to Reasoning Models and Inference-Time Scaling practice

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.