Match a job Paths Subjects Questions Quizzes Pricing
Advanced Open Pro

Diagnosing Length Inflation vs. Reward Hacking in a Training Run

Midway through an RLVR training run for a math-reasoning model, you notice average reasoning-trace length has grown from roughly 300 tokens to roughly 1,800 tokens over the last several thousand training steps, while the measured pass rate on the training-time verifiable reward has only improved from 61% to 64% over the same window. A held-out, independently-checked eval set shows pass rate essentially flat at 60-61% across that whole window.

  1. Is this pattern more consistent with "overthinking" (length inflation) or "reward hacking"? Justify using the specific numbers given, not just the general definitions.
  2. Propose the concrete change to the reward function you'd make in response, and explain the mechanism by which it addresses the diagnosed problem specifically.
  3. Why is the held-out eval result in this scenario doing more diagnostic work than the training-time reward number alone would?

Share this question

← Back to Training Reasoning Models: STaR, RLVR, and Reward Models practice

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.