Advanced
Open
Pro
Swapping a Rule-Based Verifier for a Dense PRM Reward — and Watching Accuracy Fall
Your math RLVR run learns slowly under a sparse, rule-based final-answer reward. To densify the signal, you train a Math-Shepherd-style PRM using 8 rollouts per step from the policy's step-0 checkpoint, then switch the GRPO reward to the mean PRM score across a trace's steps. Over the next 2,000 steps: PRM-scored reward climbs from 0.55 to 0.85; rule-verified final-answer accuracy on a held-out set falls from 61% to 52%; and traces fill up with short "Let me verify: ... confirmed." steps.
- Diagnose which two named failure modes are interacting, and explain why replacing a rule-based check with a learned reward opened the door where the rule-based check had not.
- Explain the distribution-shift mechanism precisely: what did the Math-Shepherd labels actually encode about the step-0 policy, why is the "verify... confirmed." tic a plausible exploit, and what does the choice of mean step score contribute?
- Redesign the reward so the PRM improves credit assignment without becoming the optimization target: name the combination of terms, the retraining discipline, and the monitoring you'd keep.
Share this question