Advanced
Open
Pro
GRPO's Advantage Formula When an Entire Sampled Group Gets the Same Reward
Training a math-reasoning policy with GRPO, group size G=8 per prompt. For a batch of "trivially easy" prompts, all 8 sampled completions get identical reward = 1 (every completion in the group is correct). For a batch of "way too hard" prompts, all 8 completions get identical reward = 0 (every completion is wrong).
- Walk through what the group-relative advantage formula
advantage_i = (reward_i - mean(group)) / std(group)computes for both of these groups specifically, and name the practical problem this causes. - This is a symptom of something larger than that day's specific batch — name the general shape of the problem it points to.
- Propose a concrete mitigation, referencing what property of a prompt actually makes a group produce a usable training signal under this design.
Share this question