Match a job Paths Subjects Questions Quizzes Pricing
Advanced Open Pro

A Math Policy Learns to Emit Two Boxed Answers: Exploiting a Loose Answer Extractor

You're running RLVR with GRPO (G=8) on a math dataset. The rule-based verifier extracts the answer with a regex over every \boxed{...} in the trace and awards reward 1 if any boxed value matches the ground truth. After several thousand steps, traces increasingly end with a pattern like "So the answer is \boxed{17}. Alternatively, it could be \boxed{19}." Training-time reward has risen from 58% to 71%; a held-out set scored under a strict one-committed-answer protocol is flat at 57%.

  1. Classify this precisely as reward hacking, a verifier coverage limit, or length inflation — explain the mechanism by which the reward as specified permits it, and name which of the subject's own examples it matches.
  2. Trace, using the group-relative advantage formula, how the hedging pattern propagates once it appears in a single completion: work through a group where one hedged completion earns reward 1 and the other seven honest completions earn 0.
  3. Propose the fix to the verifier. Explain why changing the rule to "only score the last boxed answer" is insufficient on its own, and name the audit you'd keep running after the fix.

Share this question

← Back to Training Reasoning Models: STaR, RLVR, and Reward Models practice

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.