Match a job Paths Subjects Questions Quizzes Pricing
Advanced Open Pro

GRPO's Advantage Formula When an Entire Sampled Group Gets the Same Reward

Training a math-reasoning policy with GRPO, group size G=8 per prompt. For a batch of "trivially easy" prompts, all 8 sampled completions get identical reward = 1 (every completion in the group is correct). For a batch of "way too hard" prompts, all 8 completions get identical reward = 0 (every completion is wrong).

  1. Walk through what the group-relative advantage formula advantage_i = (reward_i - mean(group)) / std(group) computes for both of these groups specifically, and name the practical problem this causes.
  2. This is a symptom of something larger than that day's specific batch — name the general shape of the problem it points to.
  3. Propose a concrete mitigation, referencing what property of a prompt actually makes a group produce a usable training signal under this design.

Share this question

← Back to Training Reasoning Models: STaR, RLVR, and Reward Models practice

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.