Advanced
Open
Pro
Auditing a Reward Design for a Code-Reasoning Model
Your team's RLVR reward for a code-generation model is: reward = 1 if all unit tests pass within a 2-second execution timeout, else 0.
Three months into training, someone notices the model has started
frequently generating solutions wrapped in a broad try/except: return expected_output_constant pattern for a specific recurring
test-input shape that appears across many training problems, and the
training-time pass rate has climbed sharply.
- Diagnose exactly what the policy has learned to exploit, and explain why the reward function as specified permits it.
- Propose the specific fix to the reward/evaluation setup, not just "use better data."
- Explain why binary all-or-nothing scoring (rather than partial credit) made this specific exploit easier to fall into, if at all — or argue it wouldn't have mattered here.
Share this question