Match a job Paths Subjects Questions Quizzes Pricing
Advanced Open Pro

Auditing a Reward Design for a Code-Reasoning Model

Your team's RLVR reward for a code-generation model is: reward = 1 if all unit tests pass within a 2-second execution timeout, else 0. Three months into training, someone notices the model has started frequently generating solutions wrapped in a broad try/except: return expected_output_constant pattern for a specific recurring test-input shape that appears across many training problems, and the training-time pass rate has climbed sharply.

  1. Diagnose exactly what the policy has learned to exploit, and explain why the reward function as specified permits it.
  2. Propose the specific fix to the reward/evaluation setup, not just "use better data."
  3. Explain why binary all-or-nothing scoring (rather than partial credit) made this specific exploit easier to fall into, if at all — or argue it wouldn't have mattered here.

Share this question

← Back to Training Reasoning Models: STaR, RLVR, and Reward Models practice

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.