Match a job Paths Subjects Questions Quizzes Pricing
Advanced Open Pro

Designing Staged Partial Credit for a Long, Multi-Function Code-Reasoning Task

Extending this subject's code-reasoning worked example: instead of a single function with 12 unit tests, you're training a model to generate an entire multi-file module (several functions, some depending on others) against a large integration test suite, and a correct solution requires many things to go right simultaneously — the code has to parse, every function has to type-check, and only then can any integration test possibly pass. Under the worked example's "fraction of tests passed" scheme applied naively here, the model's reward is 0 for the vast majority of early attempts, because a single syntax error anywhere in the module means zero tests can even run.

  1. Diagnose precisely why "fraction of tests passed" alone under-serves credit assignment specifically for this multi-file, staged-dependency task, distinct from the single-function case the worked example covers.
  2. Design a staged reward scheme for this task, naming the checkpoints and how they'd be weighted or sequenced, grounded in the subject's own Math-Shepherd-to-code analogy.
  3. Name one new reward-hacking risk this staged design introduces that the simple pass-fraction scheme didn't have, and propose a mitigation.

Share this question

← Back to Training Reasoning Models: STaR, RLVR, and Reward Models practice

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.