Advanced
Open
Pro
Would GRPO's Group-Relative Baseline Work for a Robotics RL Task?
A colleague working on a warehouse-robotics picking policy (the same
kind of task covered in ppo-and-modern-policy-optimization's
opening scenario) proposes switching from PPO with GAE to GRPO,
arguing "GRPO is strictly simpler since it drops the value network
entirely — why would we keep paying for one?"
- Would dropping the value network in favor of a group-relative baseline transfer cleanly to this robotics setting? Identify the specific structural difference between LLM-reasoning RL and this robotics task that determines the answer.
- If the team tried GRPO here anyway, what would you expect to go wrong or become harder, specifically?
- State the general rule for when a group-relative baseline is a good substitute for a learned value function, rather than treating "no value network" as an unconditional simplification.
Share this question