Practice — PPO & Modern Policy Optimization (6 questions)
Advanced
Open
Free
Diagnose and Fix a Policy Collapse During Vanilla Policy Gradient Training Permalink →
A team is training a warehouse picking-arm policy with REINFORCE plus a learned value baseline (no clipping, no trust region). Training success rate climbs steadily from 40% to 81% over 200 iterations, then on iteration 201 it drops to 6% and never recovers, even though the reward function, environment, and code did not change.
- Explain, mechanistically, why a single large policy-gradient update can cause this kind of collapse, and why it is qualitatively more dangerous than an equivalently "bad" step in supervised learning.
- Describe how TRPO's formulation would have prevented this specific failure, including what quantity it constrains and why that quantity (rather than a parameter-space step size) is the right one to bound.
- The team wants a fix that doesn't require implementing conjugate gradient and Fisher-vector products. Describe PPO's clipped surrogate objective and explain precisely why it removes the incentive for the kind of runaway update that happened here.
Share this question
Advanced
Open
Pro