Advanced
Open
Pro
Diagnosing Length Inflation vs. Reward Hacking in a Training Run
Midway through an RLVR training run for a math-reasoning model, you notice average reasoning-trace length has grown from roughly 300 tokens to roughly 1,800 tokens over the last several thousand training steps, while the measured pass rate on the training-time verifiable reward has only improved from 61% to 64% over the same window. A held-out, independently-checked eval set shows pass rate essentially flat at 60-61% across that whole window.
- Is this pattern more consistent with "overthinking" (length inflation) or "reward hacking"? Justify using the specific numbers given, not just the general definitions.
- Propose the concrete change to the reward function you'd make in response, and explain the mechanism by which it addresses the diagnosed problem specifically.
- Why is the held-out eval result in this scenario doing more diagnostic work than the training-time reward number alone would?
Share this question