Match a job Paths Subjects Questions Quizzes Pricing
Advanced Open Pro

Sizing a Compute-Matched Comparison for a Code-Generation Feature

You're deciding between two ways to spend a fixed inference-compute budget for a code-generation feature that has unit tests available as a cheap, reliable correctness check: (a) one call to a large reasoning model per request, or (b) N calls to a smaller model per request, each scored against the unit tests, keeping the first attempt that passes all tests (or the one passing the most, if none pass all).

  1. Using this subject's FLOPs-scale-with-parameters approximation, if the large model is 6x the parameter count of the small model, roughly how many small-model calls does option (b) get for the same compute as one large-model call?
  2. Explain why the unit-test check in this scenario is a much stronger selection signal than a general-purpose learned verifier used elsewhere in this subject's worked example. What property of unit tests gives it that strength?
  3. Despite that strength, name one way this approach can still fail to reflect true code correctness, and how you'd catch it before shipping.

Share this question

← Back to Reasoning Models and Inference-Time Scaling practice

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.