Advanced
Open
Pro
Sizing a Compute-Matched Comparison for a Code-Generation Feature
You're deciding between two ways to spend a fixed inference-compute budget for a code-generation feature that has unit tests available as a cheap, reliable correctness check: (a) one call to a large reasoning model per request, or (b) N calls to a smaller model per request, each scored against the unit tests, keeping the first attempt that passes all tests (or the one passing the most, if none pass all).
- Using this subject's FLOPs-scale-with-parameters approximation, if the large model is 6x the parameter count of the small model, roughly how many small-model calls does option (b) get for the same compute as one large-model call?
- Explain why the unit-test check in this scenario is a much stronger selection signal than a general-purpose learned verifier used elsewhere in this subject's worked example. What property of unit tests gives it that strength?
- Despite that strength, name one way this approach can still fail to reflect true code correctness, and how you'd catch it before shipping.
Share this question