Diagnose and Fix a Policy Collapse During Vanilla Policy Gradient Training
A team is training a warehouse picking-arm policy with REINFORCE plus a learned value baseline (no clipping, no trust region). Training success rate climbs steadily from 40% to 81% over 200 iterations, then on iteration 201 it drops to 6% and never recovers, even though the reward function, environment, and code did not change.
- Explain, mechanistically, why a single large policy-gradient update can cause this kind of collapse, and why it is qualitatively more dangerous than an equivalently "bad" step in supervised learning.
- Describe how TRPO's formulation would have prevented this specific failure, including what quantity it constrains and why that quantity (rather than a parameter-space step size) is the right one to bound.
- The team wants a fix that doesn't require implementing conjugate gradient and Fisher-vector products. Describe PPO's clipped surrogate objective and explain precisely why it removes the incentive for the kind of runaway update that happened here.
1. Why the collapse happens
Vanilla policy gradient takes a step \theta \leftarrow \theta + \alpha \nabla_\theta J(\theta) whose parameter-space size is controlled by \alpha, but what actually matters for stability is how far the output distribution moves — and there is no fixed relationship between the two. In a region of parameter space where the policy is close to deterministic (a near-saturated softmax, or a squashed continuous-action mean near the arm's workspace boundary), a small parameter change can produce a large change in behavior. On iteration 201, a minibatch with unusually large, consistent advantage estimates likely produced a gradient that pushed the policy's grasp point far enough that it started missing the box entirely.
This is more dangerous than supervised learning specifically because the policy generates its own training data. In supervised learning, a bad step still leaves you with the same fixed dataset to recover on the next step. Here, once the policy has moved into a region where it grasps at nothing, every subsequent rollout is collected under that broken policy, so the next gradient is computed from uninformative or actively misleading data — there is no fixed ground truth to fall back on, and the failure is close to self-reinforcing rather than self-correcting.
2. How TRPO prevents it
TRPO reformulates the update as: maximize the importance-weighted surrogate objective \mathbb{E}[r_t(\theta) A_t] subject to a hard constraint that the KL divergence between the old and new policy's action distributions stays under a budget \delta. KL divergence is the right quantity to bound because it directly measures how much the distribution over actions has changed, which is exactly the quantity responsible for the collapse — not how far \theta moved in Euclidean terms. TRPO approximates this constraint using the Fisher information matrix (the local curvature of KL), computes a natural gradient direction via conjugate gradient, and line-searches for the largest step that both improves the surrogate and stays inside the KL budget. Under this constraint, the update on iteration 201 would have been capped well before the policy's behavior changed enough to start missing the box — the whole point of the trust region is that step size adapts to how sensitive the policy's behavior is at the current point, not to a fixed learning rate.
3. PPO's clipped surrogate as a cheaper fix
PPO drops the explicit KL constraint and instead reshapes the objective itself: $L^{\text{CLIP}}(\theta) = \mathbb{E}_t[\min(r_t(\theta) A_t,\ \text{clip}(r_t(\theta), 1-\epsilon, 1+\epsilon) A_t)]$ For a state with positive advantage (the arm's grasp was good), once the ratio r_t(\theta) = \pi_\theta(a|s)/\pi_{\theta_\text{old}}(a|s) exceeds 1+\epsilon, the clipped term becomes the minimum and further increasing r_t no longer increases the objective — the gradient contribution from that sample vanishes. This means a batch with unusually large, consistent positive advantages (exactly what triggered the collapse) cannot keep pushing the policy arbitrarily far in one update the way the unclipped surrogate could; each sample's contribution to how far the ratio can profitably move is capped at roughly \pm\epsilon around 1. Critically, the clip does not block correction — if a later epoch's advantage sign would pull an already-drifted ratio back toward 1, the unclipped term is the minimum and the gradient still flows. This gives PPO the same qualitative protection TRPO's KL constraint provides — no single update rewards moving arbitrarily far from the old policy — using only a min and a clip inside an ordinary SGD loss, with no second-order machinery required.
Share this question