Advanced
Open
Pro
GRPO Advantages Under Partial Credit: When 11/12 Tests Passed Gets Pushed Down as Hard as 0/12
You're training a code-reasoning policy with GRPO, group size G=4,
using the worked example's partial-credit reward: reward = fraction of the 12 unit tests passed. Use the subject's advantage formula
advantage_i = (reward_i - mean(group)) / std(group) with the
population standard deviation.
- Group A's four completions pass 12/12, 9/12, 9/12, and 0/12 tests. Compute all four advantages. Then recompute them as if the reward were the binary "all tests pass or 0" scheme instead. What changes for the two 9/12 completions, and why does that change matter for credit assignment?
- Group B's four completions pass 12/12, 11/12, 11/12, and 11/12. Compute the advantages and compare them to the binary Group A result from part 1. Name the property of the formula that explains what you see, and say why it is a concern.
- The subject says partial credit preserves "a gradient of information" that all-or-nothing throws away, yet part 2 seems to show the group normalization erasing reward-scale information. Reconcile the two, and propose one concrete mitigation.
Share this question