Match a job Paths Subjects Questions Quizzes Pricing
Advanced Open Pro

GRPO Advantages Under Partial Credit: When 11/12 Tests Passed Gets Pushed Down as Hard as 0/12

You're training a code-reasoning policy with GRPO, group size G=4, using the worked example's partial-credit reward: reward = fraction of the 12 unit tests passed. Use the subject's advantage formula advantage_i = (reward_i - mean(group)) / std(group) with the population standard deviation.

  1. Group A's four completions pass 12/12, 9/12, 9/12, and 0/12 tests. Compute all four advantages. Then recompute them as if the reward were the binary "all tests pass or 0" scheme instead. What changes for the two 9/12 completions, and why does that change matter for credit assignment?
  2. Group B's four completions pass 12/12, 11/12, 11/12, and 11/12. Compute the advantages and compare them to the binary Group A result from part 1. Name the property of the formula that explains what you see, and say why it is a concern.
  3. The subject says partial credit preserves "a gradient of information" that all-or-nothing throws away, yet part 2 seems to show the group normalization erasing reward-scale information. Reconcile the two, and propose one concrete mitigation.

Share this question

← Back to Training Reasoning Models: STaR, RLVR, and Reward Models practice

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.