Match a job Paths Subjects Questions Quizzes Pricing
Overview Read Practice

Practice — PPO & Modern Policy Optimization (6 questions)

Pro content

Sign up free, then start a 14-day Pro trial — no card needed.

Advanced Open Free

Diagnose and Fix a Policy Collapse During Vanilla Policy Gradient Training Permalink →

A team is training a warehouse picking-arm policy with REINFORCE plus a learned value baseline (no clipping, no trust region). Training success rate climbs steadily from 40% to 81% over 200 iterations, then on iteration 201 it drops to 6% and never recovers, even though the reward function, environment, and code did not change.

  1. Explain, mechanistically, why a single large policy-gradient update can cause this kind of collapse, and why it is qualitatively more dangerous than an equivalently "bad" step in supervised learning.
  2. Describe how TRPO's formulation would have prevented this specific failure, including what quantity it constrains and why that quantity (rather than a parameter-space step size) is the right one to bound.
  3. The team wants a fix that doesn't require implementing conjugate gradient and Fisher-vector products. Describe PPO's clipped surrogate objective and explain precisely why it removes the incentive for the kind of runaway update that happened here.

Share this question

Advanced Open Pro

Compute the Clipped Surrogate Objective for a Batch of Samples

Unlock this question →
Advanced Open Pro

Choose a GAE Lambda for Two Different Reward Structures

Unlock this question →
Advanced Open Pro

Debug an Unstable PPO Run from Its Implementation Details

Unlock this question →
Advanced Open Pro

Map PPO's Components onto an RLHF Fine-Tuning Pipeline

Unlock this question →
Advanced Open Pro

Justify Choosing PPO Over TRPO for a Production RL System

Unlock this question →

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.